Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
Researchers have introduced ThinkARM (Anatomy of Reasoning in Models), a new framework designed to analyze the cognitive structure of large language models (LLMs) during mathematical problem-solving. Addressing the limitations of surface-level token statistics, the study adopts Schoenfeld's Episode Theory to abstract reasoning traces into functional steps such as Analysis, Exploration, Implementation, and Verification. Applied to diverse models, this approach reveals reproducible thinking dynamics and distinct structural differences between reasoning-capable and non-reasoning models. Key findings from diagnostic case studies indicate that exploration serves as a critical branching step linked to correctness. Furthermore, efficiency-oriented methods were found to selectively suppress evaluative feedback steps rather than uniformly shortening responses. This research demonstrates that episode-level representations make reasoning processes explicit, enabling a systematic analysis of how reasoning is structured, stabilized, and altered in modern AI systems. The paper, authored by Ming Li and colleagues, was published on arXiv, contributing to the fields of computer science and artificial intelligence by providing deeper insights into the internal mechanics of LLM reasoning.
Wire timeline
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
Researchers have introduced ThinkARM (Anatomy of Reasoning in Models), a new framework designed to analyze the cognitive structure of large language models (LLMs) during mathematical problem-solving. Addressing the limitations of surface-level token statistics, the study adopts Schoenfeld's Episode Theory to abstract reasoning traces into functional steps such as Analysis, Exploration, Implementation, and Verification. Applied to diverse models, this approach reveals reproducible thinking dynamics and distinct structural differences between reasoning-capable and non-reasoning models. Key findings from diagnostic case studies indicate that exploration serves as a critical branching step linked to correctness. Furthermore, efficiency-oriented methods were found to selectively suppress evaluative feedback steps rather than uniformly shortening responses. This research demonstrates that episode-level representations make reasoning processes explicit, enabling a systematic analysis of how reasoning is structured, stabilized, and altered in modern AI systems. The paper, authored by Ming Li and colleagues, was published on arXiv, contributing to the fields of computer science and artificial intelligence by providing deeper insights into the internal mechanics of LLM reasoning.
cs.AI updates on arXiv.org