GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
GigaSpeechBench is a new benchmark for automatic speech recognition (ASR) that includes 680 hours of human-annotated speech spanning low-resource languages, dialects, accents, domain-specific terminology, and age variation. The benchmark covers over a dozen Middle Eastern and Southeast Asian languages, multiple Chinese dialects, English accents, and specialized domains, with human-annotated translations for 11 languages. Evaluations of leading ASR models and commercial APIs show significant performance drops in these challenging real-world scenarios, revealing major blind spots in current ASR evaluation.
Why it matters: This benchmark exposes critical gaps in ASR robustness for over a billion underrepresented speakers, encouraging the development of more inclusive and realistic evaluation standards.
Full story at: arXiv Audio and Speech Processing ↗