AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
Researchers have introduced AdaRubric, a novel framework designed to improve the evaluation of Large Language Model (LLM) agents by generating task-specific rubrics dynamically. Unlike traditional static evaluation methods that apply fixed criteria regardless of context, AdaRubric adapts its assessment dimensions based on specific task descriptions, such as prioritizing correctness for code-debugging over fluency. The system evaluates agent trajectories step-by-step using confidence-weighted scoring and produces dense reward signals for preference learning. It employs three filtering strategies, including the new DimensionAwareFilter, to ensure high-quality data for Direct Preference Optimization (DPO). Experimental results on benchmarks like WebArena and ToolBench demonstrate a Pearson correlation of 0.79 with human evaluations, significantly outperforming existing baselines. Furthermore, DPO models trained on AdaRubric-generated data show a 6.8-8.5% improvement in task success rates. The framework also exhibits strong generalization capabilities, extending effectively to unseen domains and multimodal agents without modification, offering a robust solution for reliable LLM agent assessment.
Wire timeline
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
Researchers have introduced AdaRubric, a novel framework designed to improve the evaluation of Large Language Model (LLM) agents by generating task-specific rubrics dynamically. Unlike traditional static evaluation methods that apply fixed criteria regardless of context, AdaRubric adapts its assessment dimensions based on specific task descriptions, such as prioritizing correctness for code-debugging over fluency. The system evaluates agent trajectories step-by-step using confidence-weighted scoring and produces dense reward signals for preference learning. It employs three filtering strategies, including the new DimensionAwareFilter, to ensure high-quality data for Direct Preference Optimization (DPO). Experimental results on benchmarks like WebArena and ToolBench demonstrate a Pearson correlation of 0.79 with human evaluations, significantly outperforming existing baselines. Furthermore, DPO models trained on AdaRubric-generated data show a 6.8-8.5% improvement in task success rates. The framework also exhibits strong generalization capabilities, extending effectively to unseen domains and multimodal agents without modification, offering a robust solution for reliable LLM agent assessment.
cs.AI updates on arXiv.org