Claw-Eval: A New Framework for Trustworthy Autonomous Agent Evaluation
Researchers have introduced Claw-Eval, a comprehensive evaluation suite designed to address critical gaps in assessing autonomous agents powered by large language models. Current benchmarks often suffer from opaque grading, insufficient safety checks, and limited modality coverage. Claw-Eval features 300 human-verified tasks across nine categories, including service orchestration and multimodal interaction. It employs trajectory-aware grading using execution traces, audit logs, and environment snapshots to generate fine-grained rubrics. The framework evaluates completion, safety, and robustness, utilizing metrics like Pass@k and Pass^k to distinguish genuine capability from luck. Experiments on 14 frontier models reveal that traditional opaque evaluations miss significant safety violations and robustness failures. Furthermore, the study highlights that model capability does not guarantee consistency, with performance dropping notably under error injection. The results emphasize the multi-dimensional nature of agent capabilities, suggesting that heterogeneous evaluation is essential for developing reliably deployable autonomous systems in real-world software environments.
Wire timeline
Claw-Eval: A New Framework for Trustworthy Autonomous Agent Evaluation
Researchers have introduced Claw-Eval, a comprehensive evaluation suite designed to address critical gaps in assessing autonomous agents powered by large language models. Current benchmarks often suffer from opaque grading, insufficient safety checks, and limited modality coverage. Claw-Eval features 300 human-verified tasks across nine categories, including service orchestration and multimodal interaction. It employs trajectory-aware grading using execution traces, audit logs, and environment snapshots to generate fine-grained rubrics. The framework evaluates completion, safety, and robustness, utilizing metrics like Pass@k and Pass^k to distinguish genuine capability from luck. Experiments on 14 frontier models reveal that traditional opaque evaluations miss significant safety violations and robustness failures. Furthermore, the study highlights that model capability does not guarantee consistency, with performance dropping notably under error injection. The results emphasize the multi-dimensional nature of agent capabilities, suggesting that heterogeneous evaluation is essential for developing reliably deployable autonomous systems in real-world software environments.
cs.AI updates on arXiv.org