Medical AI Safety Ratings Vary with Judge Choice and Reveal Same-Provider Bias
A new preprint stress-tests leading medical AI models in open-ended clinical conversations with missing information, revealing that the choice of LLM judge significantly alters apparent safety ratings. The study finds only moderate agreement between judges and identifies a same-provider bias, where models appear safer when evaluated by their own provider's judge. LLM judges are also more lenient than human clinicians, indicating that observed safety gaps are due to calibration differences rather than knowledge deficits.
Why it matters: This work highlights that current evaluation practices for medical AI may overstate safety due to evaluator bias, underscoring the need for independent, human-anchored assessments before clinical deployment.
Full story at: arXiv AI/ML ↗