Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes
A new research paper investigates whether linearly decodable failure signals in Large Language Model (LLM) hidden states can be leveraged to correct errors, specifically focusing on a phenomenon termed Overthinking (OT). This behavioral regime occurs in medical question-answering tasks where models answer correctly under resampling but fail during extended chain-of-thought reasoning. The study finds that while OT is linearly decodable with 71.6% balanced accuracy, fixed linear steering methods fail to correct these failures across various configurations and architectures, including Qwen2.5-7B. Evidence suggests this limitation stems from representational entanglement, where the failure direction significantly overlaps with task-critical computations. Attempts to steer or erase these representations often damage overall model accuracy. However, the research highlights a positive application: the same decoding probe enables selective abstention, outperforming uncertainty baselines in reliability estimation. This indicates that while fixed linear steering cannot exploit decodable failure structures for correction, they remain valuable for post-generation reliability assessment in AI systems.
Wire timeline
Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes
A new research paper investigates whether linearly decodable failure signals in Large Language Model (LLM) hidden states can be leveraged to correct errors, specifically focusing on a phenomenon termed Overthinking (OT). This behavioral regime occurs in medical question-answering tasks where models answer correctly under resampling but fail during extended chain-of-thought reasoning. The study finds that while OT is linearly decodable with 71.6% balanced accuracy, fixed linear steering methods fail to correct these failures across various configurations and architectures, including Qwen2.5-7B. Evidence suggests this limitation stems from representational entanglement, where the failure direction significantly overlaps with task-critical computations. Attempts to steer or erase these representations often damage overall model accuracy. However, the research highlights a positive application: the same decoding probe enables selective abstention, outperforming uncertainty baselines in reliability estimation. This indicates that while fixed linear steering cannot exploit decodable failure structures for correction, they remain valuable for post-generation reliability assessment in AI systems.
cs.AI updates on arXiv.org