← Back to brief
Policy & SafetyReportedThe Decoder

AI Chatbots Reading X-Rays Can Be Dangerously Confident Even When They're Wrong

The RadLE 2.0 benchmark evaluates whether AI models in radiology can recognize when to defer diagnoses to human radiologists. Many AI models still make incorrect findings with high confidence, while human radiologists continue to outperform them. The study highlights the need for AI systems to learn when to abstain from making diagnoses before they can be used autonomously.

Why it matters: This research highlights a critical safety gap in medical AI: overconfident errors could lead to misdiagnosis, emphasizing the need for models that know their limits.

Full story at: The Decoder