← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Study Finds Self-Judgement Confounds Correctness Probes in Language Models

Researchers have found that hidden-state probes intended to measure output correctness in language models often capture the model's own self-judgement rather than objective correctness. By creating cases where the model's self-judgement and objective correctness disagree, they show that the self-judgement signal transfers reliably across tasks, while the objective-correctness signal does not. This indicates that transferability of probe signals does not guarantee they reflect objective truth.

Why it matters: This challenges the reliability of correctness probes for interpretability and safety, as they may conflate a model's self-assessment with actual correctness.

Full story at: arXiv Computation and Language