PSU and Duke Researchers Introduce Automated Failure Attribution for Multi-Agent Systems
Researchers from Penn State University and Duke University, in collaboration with Google DeepMind and other institutions, have introduced a novel research framework called 'Automated Failure Attribution' for Large Language Model (LLM) Multi-Agent systems. Addressing the critical challenge of diagnosing failures in complex autonomous collaborations, the team formalized the task of identifying which agent caused a failure and at what specific step. To support this, they constructed 'Who&When,' the first benchmark dataset containing 127 diverse failure logs with fine-grained human annotations detailing the responsible agent, the error timing, and the cause. The study also evaluated three automated attribution methods: All-at-Once, Step-by-Step, and Binary Search, highlighting trade-offs between cost and precision. This groundbreaking work, accepted as a Spotlight presentation at ICML 2025, aims to transform debugging from a labor-intensive manual process into a quantifiable, automated procedure. By providing open-source code and datasets, the researchers seek to enhance the reliability and iteration speed of LLM-driven multi-agent systems, addressing a significant bottleneck in current AI development workflows.
Wire timeline
PSU and Duke Researchers Introduce Automated Failure Attribution for Multi-Agent Systems
Researchers from Penn State University and Duke University, in collaboration with Google DeepMind and other institutions, have introduced a novel research framework called 'Automated Failure Attribution' for Large Language Model (LLM) Multi-Agent systems. Addressing the critical challenge of diagnosing failures in complex autonomous collaborations, the team formalized the task of identifying which agent caused a failure and at what specific step. To support this, they constructed 'Who&When,' the first benchmark dataset containing 127 diverse failure logs with fine-grained human annotations detailing the responsible agent, the error timing, and the cause. The study also evaluated three automated attribution methods: All-at-Once, Step-by-Step, and Binary Search, highlighting trade-offs between cost and precision. This groundbreaking work, accepted as a Spotlight presentation at ICML 2025, aims to transform debugging from a labor-intensive manual process into a quantifiable, automated procedure. By providing open-source code and datasets, the researchers seek to enhance the reliability and iteration speed of LLM-driven multi-agent systems, addressing a significant bottleneck in current AI development workflows.
Synced