← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Chain-of-Thought Unfaithfulness: Detection Methods Fail Most Where Models Are Wrong

A new arXiv preprint finds that most instances of unfaithful chain-of-thought (CoT) reasoning in language models occur when the model's answer is incorrect, but current behavioral detection methods are ineffective at identifying these cases. While detection methods show moderate success on correct answers, they perform no better than chance on incorrect ones, and a commonly used metric even anti-correlates with human judgments. This suggests that existing tools for auditing AI reasoning may miss the most critical failures.

Why it matters: The findings highlight a major limitation in current approaches to auditing AI reasoning, as they are least effective precisely when models make mistakes.

Full story at: arXiv Computation and Language