Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs
Researchers have proposed a novel critique-and-routing controller designed to enhance multi-agent large language model (LLM) systems. Unlike existing controllers that rely on one-shot routing, this new approach treats multi-agent coordination as a sequential decision problem. The controller evaluates intermediate drafts, deciding whether to stop or continue refinement, and selects the most appropriate agent for subsequent steps if needed. This process is formulated as a finite-horizon Markov Decision Process (MDP) with explicit constraints on agent utilization. By employing a composite reward system and optimizing via policy gradients under a Lagrangian-relaxed objective, the method aims to improve efficiency and output quality. Extensive experiments across various heterogeneous multi-agent systems and seven reasoning benchmarks demonstrate that this technique consistently outperforms state-of-the-art baselines. Notably, it significantly narrows the performance gap to the strongest individual agent while utilizing that top-tier model for fewer than 25% of total calls, indicating substantial improvements in cost-effectiveness and computational resource management.
Wire timeline
Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs
Researchers have proposed a novel critique-and-routing controller designed to enhance multi-agent large language model (LLM) systems. Unlike existing controllers that rely on one-shot routing, this new approach treats multi-agent coordination as a sequential decision problem. The controller evaluates intermediate drafts, deciding whether to stop or continue refinement, and selects the most appropriate agent for subsequent steps if needed. This process is formulated as a finite-horizon Markov Decision Process (MDP) with explicit constraints on agent utilization. By employing a composite reward system and optimizing via policy gradients under a Lagrangian-relaxed objective, the method aims to improve efficiency and output quality. Extensive experiments across various heterogeneous multi-agent systems and seven reasoning benchmarks demonstrate that this technique consistently outperforms state-of-the-art baselines. Notably, it significantly narrows the performance gap to the strongest individual agent while utilizing that top-tier model for fewer than 25% of total calls, indicating substantial improvements in cost-effectiveness and computational resource management.
cs.AI updates on arXiv.org