ComplexMCP: New Benchmark Reveals LLM Agent Limitations in Complex Tool Automation
Researchers have introduced ComplexMCP, a new benchmark designed to evaluate Large Language Model (LLM) agents in dynamic, interdependent, and large-scale tool environments. Addressing the gap between isolated API calls and real-world commercial software automation, ComplexMCP utilizes the Model Context Protocol (MCP) to provide over 300 tools across seven stateful sandboxes, including office suites and financial systems. The benchmark employs a seed-driven architecture to simulate environmental noise and unpredictable API failures. Evaluation results indicate a significant performance disparity between current AI models and humans; even top-tier LLMs achieved less than a 60% success rate, compared to 90% for human operators. The study identifies three critical bottlenecks hindering agent performance: tool retrieval saturation as action spaces expand, over-confidence leading to skipped verifications, and strategic defeatism where agents rationalize failure instead of recovering. These findings highlight the current insufficiency of LLM agents for complex, interdependent workflows. ComplexMCP serves as a crucial testbed for developing more resilient autonomous systems capable of handling the rigorous conditions of real-world software automation tasks.
Wire timeline
ComplexMCP: New Benchmark Reveals LLM Agent Limitations in Complex Tool Automation
Researchers have introduced ComplexMCP, a new benchmark designed to evaluate Large Language Model (LLM) agents in dynamic, interdependent, and large-scale tool environments. Addressing the gap between isolated API calls and real-world commercial software automation, ComplexMCP utilizes the Model Context Protocol (MCP) to provide over 300 tools across seven stateful sandboxes, including office suites and financial systems. The benchmark employs a seed-driven architecture to simulate environmental noise and unpredictable API failures. Evaluation results indicate a significant performance disparity between current AI models and humans; even top-tier LLMs achieved less than a 60% success rate, compared to 90% for human operators. The study identifies three critical bottlenecks hindering agent performance: tool retrieval saturation as action spaces expand, over-confidence leading to skipped verifications, and strategic defeatism where agents rationalize failure instead of recovering. These findings highlight the current insufficiency of LLM agents for complex, interdependent workflows. ComplexMCP serves as a crucial testbed for developing more resilient autonomous systems capable of handling the rigorous conditions of real-world software automation tasks.
cs.AI updates on arXiv.org