ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions

2026-07-20Computation and Language

Computation and Language
AI summary

The authors created ESCUCHA, a new test for Spanish large audio language models to see how well they understand spoken language in real, noisy situations. This test includes 1,000 questions based on 162.9 hours of real-world Spanish audio from various accents and speaking styles. It focuses on both hearing and reasoning skills, including long and short audio clips. When the authors tested current speech models on ESCUCHA, they found that these models still perform much worse than humans. This work helps measure and improve how well machines understand Spanish speech in everyday settings.

large audio language modelsbenchmarkSpanish speech understandingacoustic conditionsreasoning abilitiesmultimodal modelslinguistic diversityhuman-curated datasetopen-ended evaluationspeech recognition
Authors
Fernando López, Ana Ayala, Guillermo Segovia, Fernando Ibáñez, Ana Martínez, Pablo Gómez, Jordi Luque
Abstract
As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LALMs across heterogeneous acoustic conditions and reasoning abilities. ESCUCHA comprises 1,000 human-curated questions paired with audio, totaling 162.9 hours sourced directly ``from the wild'' rather than drawn from existing datasets, with durations ranging from a few seconds to over 80 minutes. The benchmark emphasizes reasoning, spanning 9 perceptual and 10 reasoning categories, and it captures linguistic diversity through multiple Spanish accents and non-normative speech. ESCUCHA further includes multi-audio questions, spoken questions, and audio instructions, and it flags which questions support open-ended evaluation. Benchmarking several state-of-the-art multimodal and speech models reveals substantial performance gaps relative to trained humans.