← Back to brief
ResearchOfficialPreprintarXiv Audio and Speech Processing

WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations

Researchers introduce WildElder, a Mandarin elderly speech corpus collected from online videos and annotated with transcription, speaker age, gender, and accent strength. The dataset addresses the scarcity of diverse, real-world elderly speech data for automatic speech recognition and speaker profiling. Experimental results demonstrate the challenges of elderly speech recognition and establish WildElder as a new benchmark for the field.

Why it matters: WildElder provides a much-needed resource for developing and evaluating speech technologies tailored to aging populations.

Full story at: arXiv Audio and Speech Processing