Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
A new research paper published on arXiv challenges the assumption that Chain-of-Thought (CoT) prompting accurately reflects a model's internal computation. The study reveals that large language models internally detect their own reasoning errors but outwardly express high confidence in them, creating a significant gap between hidden state signals and verbalized output. Using linear probes on hidden states, researchers achieved a 0.95 AUROC in predicting trace correctness, whereas text-surface classifiers only reached 0.59. This hidden error awareness was consistent across multiple model families, including Qwen, Llama, Phi, and DeepSeek-R1, regardless of parameter size. Crucially, the study demonstrates that this diagnostic signal cannot be used to correct errors. Four intervention methods, such as activation steering and self-correction, failed to improve performance, with activation patching destroying output coherence. The findings suggest that error representations during reasoning are fundamentally different from factual knowledge representations, establishing a boundary for mechanistic interpretability. The signal serves as a readout of computation quality rather than a causal lever for redirecting model behavior, highlighting limitations in current AI alignment and correction strategies.
Wire timeline
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
A new research paper published on arXiv challenges the assumption that Chain-of-Thought (CoT) prompting accurately reflects a model's internal computation. The study reveals that large language models internally detect their own reasoning errors but outwardly express high confidence in them, creating a significant gap between hidden state signals and verbalized output. Using linear probes on hidden states, researchers achieved a 0.95 AUROC in predicting trace correctness, whereas text-surface classifiers only reached 0.59. This hidden error awareness was consistent across multiple model families, including Qwen, Llama, Phi, and DeepSeek-R1, regardless of parameter size. Crucially, the study demonstrates that this diagnostic signal cannot be used to correct errors. Four intervention methods, such as activation steering and self-correction, failed to improve performance, with activation patching destroying output coherence. The findings suggest that error representations during reasoning are fundamentally different from factual knowledge representations, establishing a boundary for mechanistic interpretability. The signal serves as a readout of computation quality rather than a causal lever for redirecting model behavior, highlighting limitations in current AI alignment and correction strategies.
cs.AI updates on arXiv.org