RACER: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
A new research paper titled "Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge" addresses the trade-offs between accuracy and computational cost when using reasoning-capable large language models (LLMs) as automated judges. The study reveals that while explicit reasoning significantly enhances judgment accuracy for complex tasks like mathematics and coding, it offers limited benefits for simpler evaluations and incurs high costs. To optimize this balance, the authors propose Robust Adaptive Cost-Efficient Routing (RACER), a framework that dynamically selects between reasoning and non-reasoning models under a fixed budget. RACER formulates routing as a constrained distributionally robust optimization problem, accounting for distribution shifts via a KL-divergence uncertainty set. The method features an efficient primal-dual algorithm with theoretical guarantees of unique optimal policy and linear convergence. Experimental results demonstrate that RACER achieves superior accuracy-cost trade-offs, particularly under distribution shift conditions. This work suggests that reasoning capabilities should be applied selectively rather than universally in LLM-as-a-Judge settings to maximize efficiency and performance.
Wire timeline
RACER: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
A new research paper titled "Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge" addresses the trade-offs between accuracy and computational cost when using reasoning-capable large language models (LLMs) as automated judges. The study reveals that while explicit reasoning significantly enhances judgment accuracy for complex tasks like mathematics and coding, it offers limited benefits for simpler evaluations and incurs high costs. To optimize this balance, the authors propose Robust Adaptive Cost-Efficient Routing (RACER), a framework that dynamically selects between reasoning and non-reasoning models under a fixed budget. RACER formulates routing as a constrained distributionally robust optimization problem, accounting for distribution shifts via a KL-divergence uncertainty set. The method features an efficient primal-dual algorithm with theoretical guarantees of unique optimal policy and linear convergence. Experimental results demonstrate that RACER achieves superior accuracy-cost trade-offs, particularly under distribution shift conditions. This work suggests that reasoning capabilities should be applied selectively rather than universally in LLM-as-a-Judge settings to maximize efficiency and performance.
cs.AI updates on arXiv.org