ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions
Researchers introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate large audio language models across diverse acoustic conditions and reasoning abilities. The benchmark features 1,000 human-curated questions paired with 162.9 hours of real-world audio, encompassing multiple Spanish accents and non-normative speech. Initial evaluations show substantial performance gaps between state-of-the-art models and trained humans.
Why it matters: This benchmark fills a critical gap in evaluating audio language models for Spanish under realistic and challenging conditions, highlighting the need for more robust and capable multilingual speech AI systems.
Full story at: arXiv Computation and Language ↗