← Back to brief
ResearchOfficialPreprintarXiv Computers and Society

IyawoBench v2.0 Benchmark Reveals Hidden Failure Modes in LLM Clinical Triage for Nigerian Primary Care

A new arXiv preprint introduces IyawoBench v2.0, a benchmark designed to evaluate large language model (LLM) performance in clinical triage for Nigerian primary care using 200 synthetic vignettes based on real patient data. The study identifies three distinct failure modes—Conservative Escalation Bias, Systematic Downgrade Bias, and Middle-Tier Instability—that are not captured by conventional safety metrics. The authors demonstrate that widely used sensitivity scores can mask significant under-triage risks, such as a 77 percentage point gap in one leading model, and argue that single-ranking benchmarks are insufficient for deployment decisions in low- and middle-income country (LMIC) settings.

Why it matters: This work exposes critical limitations in current AI safety metrics for clinical triage, highlighting the need for more nuanced evaluation frameworks to ensure safe deployment in resource-limited healthcare environments.

Full story at: arXiv Computers and Society