PSU and Duke Researchers Introduce Automated Failure Attribution for LLM Multi-Agent Systems
Researchers from Penn State University and Duke University, in collaboration with Google DeepMind and other institutions, have addressed a critical challenge in Large Language Model (LLM) Multi-Agent systems: identifying the specific agent and moment responsible for task failures. While these systems are praised for collaborative problem-solving, diagnosing failures within complex, autonomous interactions is notoriously difficult and labor-intensive. To solve this, the team introduced the novel research problem of 'Automated Failure Attribution.' They developed 'Who&When,' the first benchmark dataset designed for this task, and evaluated several automated attribution methods. This work aims to streamline system iteration and enhance the reliability of multi-agent architectures by replacing manual log analysis with efficient, automated diagnostics. The research has been accepted as a Spotlight presentation at ICML 2025, a top-tier machine learning conference. Furthermore, the researchers have made their code and dataset fully open-source, providing the global AI community with essential tools to improve the robustness and transparency of increasingly complex LLM-based collaborative systems.
Wire timeline
PSU and Duke Researchers Introduce Automated Failure Attribution for LLM Multi-Agent Systems
Researchers from Penn State University and Duke University, in collaboration with Google DeepMind and other institutions, have addressed a critical challenge in Large Language Model (LLM) Multi-Agent systems: identifying the specific agent and moment responsible for task failures. While these systems are praised for collaborative problem-solving, diagnosing failures within complex, autonomous interactions is notoriously difficult and labor-intensive. To solve this, the team introduced the novel research problem of 'Automated Failure Attribution.' They developed 'Who&When,' the first benchmark dataset designed for this task, and evaluated several automated attribution methods. This work aims to streamline system iteration and enhance the reliability of multi-agent architectures by replacing manual log analysis with efficient, automated diagnostics. The research has been accepted as a Spotlight presentation at ICML 2025, a top-tier machine learning conference. Furthermore, the researchers have made their code and dataset fully open-source, providing the global AI community with essential tools to improve the robustness and transparency of increasingly complex LLM-based collaborative systems.
Synced