Explicit Reasoning Enhances LLM Judge Accuracy, Efficiency, and Robustness
A new systematic study published on arXiv investigates the effectiveness of Large Language Models (LLMs) as automated judges in benchmarking and reward modeling. The research compares "thinking" models, which utilize explicit reasoning, against "non-thinking" models using open-source Qwen 3 variants with 0.6B, 1.7B, and 4B parameters. Evaluations on RewardBench tasks reveal that thinking models achieve approximately 10% higher accuracy with minimal computational overhead (under 2x), significantly outperforming augmentation strategies like few-shot learning that incur higher costs (>8x) for modest gains. Furthermore, the study demonstrates that thinking models exhibit greater robustness, maintaining 6% higher consistency across various bias conditions, including positional, identity, and bandwagon biases. These benefits extend to multilingual settings, confirming that explicit reasoning improves performance beyond English. The findings provide strong evidence that integrating explicit reasoning into LLMs enhances their reliability, efficiency, and robustness as automated judges, offering a superior alternative to traditional non-thinking approaches and complex augmentation techniques.
Wire timeline
Explicit Reasoning Enhances LLM Judge Accuracy, Efficiency, and Robustness
A new systematic study published on arXiv investigates the effectiveness of Large Language Models (LLMs) as automated judges in benchmarking and reward modeling. The research compares "thinking" models, which utilize explicit reasoning, against "non-thinking" models using open-source Qwen 3 variants with 0.6B, 1.7B, and 4B parameters. Evaluations on RewardBench tasks reveal that thinking models achieve approximately 10% higher accuracy with minimal computational overhead (under 2x), significantly outperforming augmentation strategies like few-shot learning that incur higher costs (>8x) for modest gains. Furthermore, the study demonstrates that thinking models exhibit greater robustness, maintaining 6% higher consistency across various bias conditions, including positional, identity, and bandwagon biases. These benefits extend to multilingual settings, confirming that explicit reasoning improves performance beyond English. The findings provide strong evidence that integrating explicit reasoning into LLMs enhances their reliability, efficiency, and robustness as automated judges, offering a superior alternative to traditional non-thinking approaches and complex augmentation techniques.
cs.AI updates on arXiv.org