GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
Researchers have introduced GR-Ben, a new benchmark designed to evaluate Process Reward Models (PRMs) beyond the traditional scope of mathematical reasoning. While PRMs show potential for improving Large Language Model (LLM) reasoning by detecting intermediate errors, existing benchmarks lack diversity. GR-Ben addresses this gap by assessing performance across science and logic domains, including nine subdomains. Extensive experiments involving 22 models revealed that current PRMs and LLMs exhibit significantly weaker error-detection capabilities in non-mathematical contexts. The study highlights a distinct divergence in strengths: PRMs struggle with identifying knowledge-based errors, whereas LLMs perform poorly in detecting computational mistakes. This research aims to foster advancements in general-domain PRMs, ultimately enhancing the robustness and reasoning accuracy of LLMs in real-world scenarios. The paper, authored by Zhouhao Sun and colleagues, was published on arXiv in May 2026, contributing to the fields of Artificial Intelligence and Computation and Language.
Wire timeline
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
Researchers have introduced GR-Ben, a new benchmark designed to evaluate Process Reward Models (PRMs) beyond the traditional scope of mathematical reasoning. While PRMs show potential for improving Large Language Model (LLM) reasoning by detecting intermediate errors, existing benchmarks lack diversity. GR-Ben addresses this gap by assessing performance across science and logic domains, including nine subdomains. Extensive experiments involving 22 models revealed that current PRMs and LLMs exhibit significantly weaker error-detection capabilities in non-mathematical contexts. The study highlights a distinct divergence in strengths: PRMs struggle with identifying knowledge-based errors, whereas LLMs perform poorly in detecting computational mistakes. This research aims to foster advancements in general-domain PRMs, ultimately enhancing the robustness and reasoning accuracy of LLMs in real-world scenarios. The paper, authored by Zhouhao Sun and colleagues, was published on arXiv in May 2026, contributing to the fields of Artificial Intelligence and Computation and Language.
cs.AI updates on arXiv.org