MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs
Researchers have introduced MAGE (Multi-Agent Graph-guided Evolution), a novel framework designed to enable self-evolving language-model agents without modifying their underlying backbone models. Addressing limitations in existing systems that rely on flat memory or implicit signals, MAGE externalizes self-knowledge into a four-subgraph co-evolutionary knowledge graph. This structure stores both teacher-written failure corrections and the learner’s successful reasoning traces, providing task-conditioned guidance to a frozen execution model. The framework utilizes task-level and skill-level routing bandits updated from a shared reward stream to optimize performance. Structural analysis highlights how append-only memory growth and bounded curriculum coverage ensure stable improvement. Evaluations across nine diverse benchmarks, including mathematical reasoning, medical multiple-choice questions, and web navigation, demonstrate that MAGE significantly outperforms prompt-based frozen-backbone baselines. Ablation studies reveal that success traces and corrective memories are complementary, with each contributing uniquely to different types of reasoning tasks. This research offers a robust solution for enhancing agent capabilities while maintaining computational efficiency by keeping the core model static.
Wire timeline
MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs
Researchers have introduced MAGE (Multi-Agent Graph-guided Evolution), a novel framework designed to enable self-evolving language-model agents without modifying their underlying backbone models. Addressing limitations in existing systems that rely on flat memory or implicit signals, MAGE externalizes self-knowledge into a four-subgraph co-evolutionary knowledge graph. This structure stores both teacher-written failure corrections and the learner’s successful reasoning traces, providing task-conditioned guidance to a frozen execution model. The framework utilizes task-level and skill-level routing bandits updated from a shared reward stream to optimize performance. Structural analysis highlights how append-only memory growth and bounded curriculum coverage ensure stable improvement. Evaluations across nine diverse benchmarks, including mathematical reasoning, medical multiple-choice questions, and web navigation, demonstrate that MAGE significantly outperforms prompt-based frozen-backbone baselines. Ablation studies reveal that success traces and corrective memories are complementary, with each contributing uniquely to different types of reasoning tasks. This research offers a robust solution for enhancing agent capabilities while maintaining computational efficiency by keeping the core model static.
cs.AI updates on arXiv.org