Global, Guideline-Grounded Evaluation Reveals Systematic Failures of XAI Methods in ECG Classification
A new preprint introduces a global, clinically grounded framework for evaluating explainable AI (XAI) methods in ECG classification. The study finds that many commonly used gradient-based XAI methods systematically fail to highlight clinically relevant regions, often focusing on signal amplitude rather than guideline-defined diagnostic features. In tests across four classifiers and 13 XAI methods, nine methods performed below chance for at least one condition, revealing inconsistent reliability. The results suggest that standard XAI approaches may misrepresent model behavior in medical contexts.
Why it matters: The findings raise concerns about the reliability of widely used XAI methods in medical AI, with implications for trust and safety in clinical decision support.
Full story at: arXiv Computers and Society ↗