TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
Researchers have introduced TMAS, a novel framework designed to enhance the reasoning capabilities of large language models through test-time compute scaling. Published on arXiv, this approach addresses limitations in existing methods by organizing inference as a collaborative process among specialized agents. TMAS utilizes hierarchical memories, including an experience bank for reusing reliable intermediate conclusions and a guideline bank to record high-level strategies, thereby preventing redundant reasoning patterns. The framework also incorporates a hybrid reward reinforcement learning scheme that balances basic reasoning preservation with enhanced experience utilization and exploration. Extensive experiments on challenging reasoning benchmarks demonstrate that TMAS achieves superior iterative scaling performance compared to current baselines. Additionally, the hybrid reward training improves both the effectiveness and stability of scaling across multiple iterations. This development represents a significant advancement in structured test-time scaling, offering a more efficient way to balance exploration and exploitation during model inference. The associated code and data have been made publicly available to support further research and replication.
Wire timeline
TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
Researchers have introduced TMAS, a novel framework designed to enhance the reasoning capabilities of large language models through test-time compute scaling. Published on arXiv, this approach addresses limitations in existing methods by organizing inference as a collaborative process among specialized agents. TMAS utilizes hierarchical memories, including an experience bank for reusing reliable intermediate conclusions and a guideline bank to record high-level strategies, thereby preventing redundant reasoning patterns. The framework also incorporates a hybrid reward reinforcement learning scheme that balances basic reasoning preservation with enhanced experience utilization and exploration. Extensive experiments on challenging reasoning benchmarks demonstrate that TMAS achieves superior iterative scaling performance compared to current baselines. Additionally, the hybrid reward training improves both the effectiveness and stability of scaling across multiple iterations. This development represents a significant advancement in structured test-time scaling, offering a more efficient way to balance exploration and exploitation during model inference. The associated code and data have been made publicly available to support further research and replication.
cs.AI updates on arXiv.org