← Back to brief
ResearchOfficialPreprintarXiv Information Retrieval

HiEviDR-Bench: Benchmarking Hierarchical Evidence Aggregation in Deep Research Tasks

A new arXiv preprint introduces HiEviDR-Bench, a benchmark designed to evaluate how well AI models aggregate and trace evidence in complex research tasks. The benchmark includes 2,000 human-validated questions with explicit evidence graphs, spanning both text-only and multimodal scenarios. Tests on 16 multimodal large language models reveal that while these systems often generate high-quality reports, they perform poorly on citation accuracy and constructing well-supported claims.

Why it matters: This work highlights a significant gap between the appearance of report quality and the underlying reasoning and evidence-tracing abilities of current AI models, which is crucial for trustworthy research automation.

Full story at: arXiv Information Retrieval