Study Reveals Format Confound in Chain-of-Thought Faithfulness Evaluations
A new research paper published on arXiv identifies a significant methodological flaw in corruption studies used to evaluate the faithfulness of Chain-of-Thought (CoT) reasoning in large language models. The study demonstrates that for chains containing explicit terminal answer statements, such as "the answer is X," corruption metrics primarily detect the location of the answer text rather than identifying computationally important reasoning steps. Experimental evidence from GSM8K and MATH benchmarks shows that removing only the final answer statement drastically reduces suffix sensitivity, while conflicting-answer experiments reveal that models systematically follow explicit answer text over preceding reasoning, particularly in models sized 3B to 14B. This format-determination effect diminishes at larger scales, converging toward zero at 32B parameters. The authors propose a new three-prerequisite protocol, including question-only controls and format characterization, to establish a minimum standard for future faithfulness studies, ensuring that evaluations accurately reflect computational processes rather than superficial formatting cues.
Wire timeline
Study Reveals Format Confound in Chain-of-Thought Faithfulness Evaluations
A new research paper published on arXiv identifies a significant methodological flaw in corruption studies used to evaluate the faithfulness of Chain-of-Thought (CoT) reasoning in large language models. The study demonstrates that for chains containing explicit terminal answer statements, such as "the answer is X," corruption metrics primarily detect the location of the answer text rather than identifying computationally important reasoning steps. Experimental evidence from GSM8K and MATH benchmarks shows that removing only the final answer statement drastically reduces suffix sensitivity, while conflicting-answer experiments reveal that models systematically follow explicit answer text over preceding reasoning, particularly in models sized 3B to 14B. This format-determination effect diminishes at larger scales, converging toward zero at 32B parameters. The authors propose a new three-prerequisite protocol, including question-only controls and format characterization, to establish a minimum standard for future faithfulness studies, ensuring that evaluations accurately reflect computational processes rather than superficial formatting cues.
cs.AI updates on arXiv.org