Correlated Agreement Blindness: A Structural Risk in Multi-Agent AI Arbitration
A new arXiv preprint identifies 'correlated agreement blindness,' a structural risk in multi-agent AI systems where improved base models tend to agree, causing disagreement-based safety monitoring to miss correlated failures. The authors introduce ARAT, a system combining diverse model types and a meta-model, which reduces under-prediction rates on network intrusion and clinical readmission datasets. The findings suggest that simply improving or diversifying models does not guarantee safer arbitration unless it leads to meaningful disagreement.
Why it matters: This work highlights a fundamental limitation in widely used disagreement-based safety mechanisms for multi-agent AI, with implications for the reliability of future agentic pipelines as models become more capable and correlated.
Full story at: arXiv Multiagent Systems ↗