Back to Research papers
Research paper index

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions

Fernando López, Ana Ayala, Guillermo Segovia, Fernando Ibáñez, Ana Martínez, Pablo Gómez, Jordi Luque

arXiv:2607.17812Published July 20, 20260 citations
  • cs.CL

Abstract

As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LALMs across heterogeneous acoustic conditions and reasoning abilities. ESCUCHA comprises 1,000 human-curated questions paired with audio, totaling 162.9 hours sourced directly ``from the wild'' rather than drawn from existing datasets, with durations ranging from a few seconds to over 80 minutes. The benchmark emphasizes reasoning, spanning 9 perceptual and 10 reasoning categories, and it captures linguistic diversity through multiple Spanish accents and non-normative speech. ESCUCHA further includes multi-audio questions, spoken questions, and audio instructions, and it flags which questions support open-ended evaluation. Benchmarking several state-of-the-art multimodal and speech models reveals substantial performance gaps relative to trained humans.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.