← Back to brief
ResearchOfficialPreprintarXiv Information Retrieval

DRNOISE Benchmark Reveals Deep Research Agents Vulnerable to Misleading Evidence

A new benchmark, DRNOISE, evaluates deep research agents on 100 tasks by introducing a single plausible but false document into search results. This intervention leads to accuracy drops of 66-88 percentage points, with agents frequently failing to complete evidence chains and often deferring to the misleading document.

Why it matters: This exposes a significant vulnerability in open-web AI agents, showing that even ordinary-looking falsehoods can seriously undermine their evidential reasoning.

Full story at: arXiv Information Retrieval