FormalRewardBench: A New Benchmark for Evaluating Reward Models in Formal Theorem Proving
Researchers have introduced FormalRewardBench, the first benchmark designed to evaluate reward models in formal theorem proving using Lean 4. Addressing the limitations of binary correctness signals in reinforcement learning, which suffer from sparse credit assignment, this new tool assesses learned reward models that evaluate proof quality beyond simple verification. The benchmark comprises 250 preference pairs, matching correct proofs with incorrect variants generated through five expert-curated error injection strategies, including forced mistakes and natural language justification. The study evaluated various models, including frontier LLMs like Claude Opus 4.5, judge LLMs, general-purpose models like Qwen2.5-72B-Instruct, and specialized provers such as DeepSeek-Prover-V2-7B. Results indicated that frontier LLMs achieved the highest performance at 59.8%, while specialized theorem provers performed worst at 24.4%, suggesting that theorem-proving capabilities do not necessarily transfer to proof evaluation. The authors publicly released FormalRewardBench to encourage further research into developing robust reward models for formal mathematics, highlighting the challenges posed by different error injection mechanisms.
Wire timeline
FormalRewardBench: A New Benchmark for Evaluating Reward Models in Formal Theorem Proving
Researchers have introduced FormalRewardBench, the first benchmark designed to evaluate reward models in formal theorem proving using Lean 4. Addressing the limitations of binary correctness signals in reinforcement learning, which suffer from sparse credit assignment, this new tool assesses learned reward models that evaluate proof quality beyond simple verification. The benchmark comprises 250 preference pairs, matching correct proofs with incorrect variants generated through five expert-curated error injection strategies, including forced mistakes and natural language justification. The study evaluated various models, including frontier LLMs like Claude Opus 4.5, judge LLMs, general-purpose models like Qwen2.5-72B-Instruct, and specialized provers such as DeepSeek-Prover-V2-7B. Results indicated that frontier LLMs achieved the highest performance at 59.8%, while specialized theorem provers performed worst at 24.4%, suggesting that theorem-proving capabilities do not necessarily transfer to proof evaluation. The authors publicly released FormalRewardBench to encourage further research into developing robust reward models for formal mathematics, highlighting the challenges posed by different error injection mechanisms.
cs.AI updates on arXiv.org