Plan-RewardBench: A New Benchmark for Trajectory-Level Reward Modeling in Agentic AI
Researchers have introduced Plan-RewardBench, a novel benchmark designed to evaluate Reward Models (RMs) within tool-integrated agentic systems. As Large Language Models evolve into autonomous agents capable of complex reasoning and tool invocation, traditional Reinforcement Learning from Human Feedback (RLHF) methods face significant challenges due to the lack of specialized assessment tools. Plan-RewardBench addresses this gap by providing a trajectory-level preference benchmark that tests how well judges distinguish between preferred and distractor agent trajectories. The benchmark covers four key task families: Safety Refusal, Tool-Irrelevance/Unavailability, Complex Planning, and Robust Error Recovery. It utilizes validated positive trajectories and hard negatives generated through multi-model rollouts and perturbations. The study benchmarks generative, discriminative, and LLM-as-Judge evaluators, revealing substantial performance degradation on long-horizon trajectories across all types. These findings highlight the urgent need for specialized training in agentic, trajectory-level reward modeling. Ultimately, this work serves as both a practical evaluation suite and a blueprint for constructing preference data for future agentic planning systems.
Wire timeline
Plan-RewardBench: A New Benchmark for Trajectory-Level Reward Modeling in Agentic AI
Researchers have introduced Plan-RewardBench, a novel benchmark designed to evaluate Reward Models (RMs) within tool-integrated agentic systems. As Large Language Models evolve into autonomous agents capable of complex reasoning and tool invocation, traditional Reinforcement Learning from Human Feedback (RLHF) methods face significant challenges due to the lack of specialized assessment tools. Plan-RewardBench addresses this gap by providing a trajectory-level preference benchmark that tests how well judges distinguish between preferred and distractor agent trajectories. The benchmark covers four key task families: Safety Refusal, Tool-Irrelevance/Unavailability, Complex Planning, and Robust Error Recovery. It utilizes validated positive trajectories and hard negatives generated through multi-model rollouts and perturbations. The study benchmarks generative, discriminative, and LLM-as-Judge evaluators, revealing substantial performance degradation on long-horizon trajectories across all types. These findings highlight the urgent need for specialized training in agentic, trajectory-level reward modeling. Ultimately, this work serves as both a practical evaluation suite and a blueprint for constructing preference data for future agentic planning systems.
cs.AI updates on arXiv.org