Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
A new research paper submitted to arXiv addresses critical reliability issues in interactive agent benchmarks. Current evaluation methods often rely on surface-level signals, such as verifying a button click rather than confirming the intended state change, leading to misleading success scores. To resolve this, the authors introduce an outcome evidence reporting layer that operates without modifying existing tasks, agents, or evaluators. This layer specifies required artifacts for verification, applies a locked checklist to assign evidence labels (Pass, Fail, or Unknown), and reports score bounds that quantify uncertainty from unknown cases. By making uncertain outcomes explicitly visible instead of hiding them within aggregate rates, the framework provides a more accurate assessment of benchmark quality. The study evaluates this approach on five public benchmarks, including ANDROIDWORLD, AGENTDOJO, APPWORLD, tau3 bench retail, and MINIWOB. The results demonstrate the ability to separate distinct failure modes, offering a robust method for improving the validity of AI agent evaluations. This work highlights the importance of reliable outcome detection alongside task design in artificial intelligence research.
Wire timeline
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
A new research paper submitted to arXiv addresses critical reliability issues in interactive agent benchmarks. Current evaluation methods often rely on surface-level signals, such as verifying a button click rather than confirming the intended state change, leading to misleading success scores. To resolve this, the authors introduce an outcome evidence reporting layer that operates without modifying existing tasks, agents, or evaluators. This layer specifies required artifacts for verification, applies a locked checklist to assign evidence labels (Pass, Fail, or Unknown), and reports score bounds that quantify uncertainty from unknown cases. By making uncertain outcomes explicitly visible instead of hiding them within aggregate rates, the framework provides a more accurate assessment of benchmark quality. The study evaluates this approach on five public benchmarks, including ANDROIDWORLD, AGENTDOJO, APPWORLD, tau3 bench retail, and MINIWOB. The results demonstrate the ability to separate distinct failure modes, offering a robust method for improving the validity of AI agent evaluations. This work highlights the importance of reliable outcome detection alongside task design in artificial intelligence research.
cs.AI updates on arXiv.org