MultiRef-Compass: Comprehensive Benchmark for Multi-Reference Audio-Video Generation
Researchers have introduced MultiRef-Compass, a benchmark designed to evaluate multi-reference-to-audio-video (MR2AV) generation systems. The benchmark consists of 350 curated samples that test capabilities such as multi-view subject preservation, multi-entity binding, and human-object-scene composition. It features 14 sub-metrics across four evaluation dimensions and integrates both automatic and model-based judging frameworks. Experiments on eight MR2AV systems demonstrate significant gaps in current model performance, highlighting the challenge of this task.
Why it matters: MultiRef-Compass addresses a critical gap in benchmarking models that must generate synchronized audio and video content from multiple references, supporting progress in advanced multimodal generation.
Full story at: arXiv Computer Vision ↗