MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs
Researchers have introduced MathConstraint, a new adaptive benchmark designed to evaluate the combinatorial reasoning capabilities of Large Language Models (LLMs). Unlike static datasets that quickly saturate, MathConstraint combines constraint satisfaction problems with rigorous solver-based verification to generate arbitrarily difficult, automatically verifiable instances. The study releases two datasets: MathConstraint-Easy and the harder MathConstraint. Performance tests on twelve frontier models reveal significant accuracy drops on the harder set, demonstrating the benchmark's resilience against rapid AI improvements. For instance, while top models achieved up to 87.6% accuracy on the easy set, this fell to between 18.5% and 66.9% on the standard set. The research also highlights the critical role of tool use; access to a sandboxed Python environment with SAT/SMT solvers roughly doubled frontier accuracy. Furthermore, reducing the tool-call budget significantly degraded performance, a sensitivity often missed by other benchmarks. The authors release the generator, dataset, and evaluation harness to support robust studies on combinatorial reasoning and tool-use behavior under tunable difficulty levels.
Wire timeline
MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs
Researchers have introduced MathConstraint, a new adaptive benchmark designed to evaluate the combinatorial reasoning capabilities of Large Language Models (LLMs). Unlike static datasets that quickly saturate, MathConstraint combines constraint satisfaction problems with rigorous solver-based verification to generate arbitrarily difficult, automatically verifiable instances. The study releases two datasets: MathConstraint-Easy and the harder MathConstraint. Performance tests on twelve frontier models reveal significant accuracy drops on the harder set, demonstrating the benchmark's resilience against rapid AI improvements. For instance, while top models achieved up to 87.6% accuracy on the easy set, this fell to between 18.5% and 66.9% on the standard set. The research also highlights the critical role of tool use; access to a sandboxed Python environment with SAT/SMT solvers roughly doubled frontier accuracy. Furthermore, reducing the tool-call budget significantly degraded performance, a sensitivity often missed by other benchmarks. The authors release the generator, dataset, and evaluation harness to support robust studies on combinatorial reasoning and tool-use behavior under tunable difficulty levels.
cs.AI updates on arXiv.org