Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
Researchers have introduced ScaleLogic, a new synthetic logical reasoning framework designed to systematically study how Reinforcement Learning (RL) training scales with task difficulty for Large Language Models (LLMs). Addressing the lack of controlled environments, ScaleLogic allows independent control over proof planning depth and logical expressiveness, ranging from simple implications to complex first-order logic. The study reveals that RL training compute follows a power law relative to reasoning depth, with the scaling exponent increasing monotonically as logical expressiveness grows. Crucially, the findings demonstrate that training on more expressive logic yields significantly larger performance gains on downstream mathematics and general reasoning benchmarks, alongside more compute-efficient transfer. This indicates that the nature of training data, not just volume, critically shapes model capabilities. Additionally, the research confirms that curriculum-based training substantially improves scaling efficiency across multiple RL methods, offering new insights into optimizing long-horizon reasoning in AI systems.
Wire timeline
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
Researchers have introduced ScaleLogic, a new synthetic logical reasoning framework designed to systematically study how Reinforcement Learning (RL) training scales with task difficulty for Large Language Models (LLMs). Addressing the lack of controlled environments, ScaleLogic allows independent control over proof planning depth and logical expressiveness, ranging from simple implications to complex first-order logic. The study reveals that RL training compute follows a power law relative to reasoning depth, with the scaling exponent increasing monotonically as logical expressiveness grows. Crucially, the findings demonstrate that training on more expressive logic yields significantly larger performance gains on downstream mathematics and general reasoning benchmarks, alongside more compute-efficient transfer. This indicates that the nature of training data, not just volume, critically shapes model capabilities. Additionally, the research confirms that curriculum-based training substantially improves scaling efficiency across multiple RL methods, offering new insights into optimizing long-horizon reasoning in AI systems.
cs.AI updates on arXiv.org