← Back to brief
ResearchOfficialPreprintarXiv Audio and Speech Processing

Echoes: A Semantically-Aligned Music Deepfake Detection Dataset

Researchers introduce Echoes, a dataset of 4,468 tracks (131 hours) spanning multiple genres and generated by ten AI music systems, designed to train and benchmark robust deepfake detectors. The dataset enforces semantic alignment between spoofed and bona fide audio to prevent shortcut learning. Cross-dataset evaluations show Echoes is the hardest in-domain dataset and that training on it yields the strongest generalization performance for deepfake detection.

Why it matters: Echoes provides a challenging and diverse benchmark that advances the robustness and generalization of AI-generated music deepfake detectors.

Full story at: arXiv Audio and Speech Processing