← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

Researchers introduce MamaBench, the first counterfactual benchmark for maternal and pediatric AI, comprising 434 expert-authored clinical narratives in 217 pairs across 371 pathologies. They propose Evidence-Anchored RAG (EA-RAG), a retrieval method that reduces the Bias Trap Rate (BTR) to 20.3% on Claude Sonnet 4.6—a 5.5 percentage point improvement—without degrading base accuracy. The study finds that standard medical benchmarks overstate LLM robustness by 16-28 percentage points, and that counterfactual robustness remains a significant challenge.

Why it matters: This work exposes critical gaps in current medical AI evaluation and introduces new tools for assessing and improving LLM robustness in clinical settings.

Full story at: arXiv Computation and Language