First Systematic Study Evaluates General-Purpose AI Agent Architectures
Researchers have released the first systematic study evaluating how different architectures impact the performance of general-purpose AI agents across diverse, unfamiliar environments. Addressing previous gaps in evaluation harnesses and benchmark compatibility, the team introduced a unifying protocol and an open evaluation framework. The study compares five agent architectures, including tool-calling and code-generation, against five backbone large language models across six benchmarks covering software engineering, customer service, and research. Key findings indicate that while backbone model choice primarily drives overall performance, architecture selection can cause performance variations of up to 12 percentage points. Notably, top general agents matched heavily customized domain-specific agents in four out of six tests. The research also identified "generality sinks" in open-weight models, where performance collapses on specific tasks, a trait absent in frontier closed-source models. This work establishes the first Open General Agent Leaderboard, providing critical insights into agent adaptability and error signatures without requiring per-domain customization.
Wire timeline
First Systematic Study Evaluates General-Purpose AI Agent Architectures
Researchers have released the first systematic study evaluating how different architectures impact the performance of general-purpose AI agents across diverse, unfamiliar environments. Addressing previous gaps in evaluation harnesses and benchmark compatibility, the team introduced a unifying protocol and an open evaluation framework. The study compares five agent architectures, including tool-calling and code-generation, against five backbone large language models across six benchmarks covering software engineering, customer service, and research. Key findings indicate that while backbone model choice primarily drives overall performance, architecture selection can cause performance variations of up to 12 percentage points. Notably, top general agents matched heavily customized domain-specific agents in four out of six tests. The research also identified "generality sinks" in open-weight models, where performance collapses on specific tasks, a trait absent in frontier closed-source models. This work establishes the first Open General Agent Leaderboard, providing critical insights into agent adaptability and error signatures without requiring per-domain customization.
cs.AI updates on arXiv.org