WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
Researchers introduce WildElder, a Mandarin elderly speech corpus collected from online videos and annotated with transcription, speaker age, gender, and accent strength. The dataset addresses the scarcity of diverse, real-world elderly speech data for automatic speech recognition and speaker profiling. Experimental results demonstrate the challenges of elderly speech recognition and establish WildElder as a new benchmark for the field.
Why it matters: WildElder provides a much-needed resource for developing and evaluating speech technologies tailored to aging populations.
Full story at: arXiv Audio and Speech Processing ↗