CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
Researchers have introduced CORE (Concept-Oriented REinforcement), a new reinforcement learning framework designed to enhance the mathematical reasoning capabilities of Large Language Models (LLMs). While current models often solve complex math problems through pattern recognition, they frequently fail to apply underlying concepts when genuine understanding is required. Existing Reinforcement Learning with Verifiable Rewards (RLVR) methods primarily reinforce final answers, lacking fine-grained conceptual signals. CORE addresses this by transforming explicit concepts into controllable supervision signals using high-quality textbook resources. The framework synthesizes concept-aligned quizzes, injects brief concept snippets during model rollouts to prime trajectories, and reinforces conceptual reasoning through trajectory replacement or standard GRPO. Empirical results demonstrate that CORE consistently outperforms vanilla and Supervised Fine-Tuning (SFT) baselines across both in-domain concept exercises and diverse out-of-domain math benchmarks. By unifying direct training on concept-aligned quizzes with concept-injected rollouts under outcome regularization, CORE effectively bridges the gap between problem-solving competence and true conceptual reasoning, offering an algorithm- and verifier-agnostic solution for improving AI mathematical logic.
Wire timeline
CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
Researchers have introduced CORE (Concept-Oriented REinforcement), a new reinforcement learning framework designed to enhance the mathematical reasoning capabilities of Large Language Models (LLMs). While current models often solve complex math problems through pattern recognition, they frequently fail to apply underlying concepts when genuine understanding is required. Existing Reinforcement Learning with Verifiable Rewards (RLVR) methods primarily reinforce final answers, lacking fine-grained conceptual signals. CORE addresses this by transforming explicit concepts into controllable supervision signals using high-quality textbook resources. The framework synthesizes concept-aligned quizzes, injects brief concept snippets during model rollouts to prime trajectories, and reinforces conceptual reasoning through trajectory replacement or standard GRPO. Empirical results demonstrate that CORE consistently outperforms vanilla and Supervised Fine-Tuning (SFT) baselines across both in-domain concept exercises and diverse out-of-domain math benchmarks. By unifying direct training on concept-aligned quizzes with concept-injected rollouts under outcome regularization, CORE effectively bridges the gap between problem-solving competence and true conceptual reasoning, offering an algorithm- and verifier-agnostic solution for improving AI mathematical logic.
cs.AI updates on arXiv.org