HiEviDR-Bench: Benchmarking Hierarchical Evidence Aggregation in Deep Research Tasks
A new arXiv preprint introduces HiEviDR-Bench, a benchmark designed to evaluate how well AI models aggregate and trace evidence in complex research tasks. The benchmark includes 2,000 human-validated questions with explicit evidence graphs, spanning both text-only and multimodal scenarios. Tests on 16 multimodal large language models reveal that while these systems often generate high-quality reports, they perform poorly on citation accuracy and constructing well-supported claims.
Why it matters: This work highlights a significant gap between the appearance of report quality and the underlying reasoning and evidence-tracing abilities of current AI models, which is crucial for trustworthy research automation.
Full story at: arXiv Information Retrieval ↗