Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
A new research paper published on arXiv introduces an automated research framework driven by specialist agents operating in a closed empirical loop. This system autonomously generates hypotheses, executes code edits, and evaluates outcomes without human intervention during the search process. The study highlights that lineage feedback allows agents to transform evaluator outcomes, such as crashes or accuracy misses, into refined program-level recipe edits. In extensive trials involving 1,197 headline runs and 600 control trials, the system demonstrated significant improvements: reducing Parameter Golf validation bits per byte by 0.81%, increasing NanoChat-D12 CORE performance by 38.7%, and decreasing CIFAR-10 Airbench96 wallclock time by 4.59%. The output is an auditable trajectory of proposals and experiments rather than just a final model. This approach represents a significant advancement in multi-agent systems and artificial intelligence, showcasing the ability of autonomous agents to write code, submit experiments, and improve public starting recipes through iterative feedback and strict architecture-domain audits.
Wire timeline
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
A new research paper published on arXiv introduces an automated research framework driven by specialist agents operating in a closed empirical loop. This system autonomously generates hypotheses, executes code edits, and evaluates outcomes without human intervention during the search process. The study highlights that lineage feedback allows agents to transform evaluator outcomes, such as crashes or accuracy misses, into refined program-level recipe edits. In extensive trials involving 1,197 headline runs and 600 control trials, the system demonstrated significant improvements: reducing Parameter Golf validation bits per byte by 0.81%, increasing NanoChat-D12 CORE performance by 38.7%, and decreasing CIFAR-10 Airbench96 wallclock time by 4.59%. The output is an auditable trajectory of proposals and experiments rather than just a final model. This approach represents a significant advancement in multi-agent systems and artificial intelligence, showcasing the ability of autonomous agents to write code, submit experiments, and improve public starting recipes through iterative feedback and strict architecture-domain audits.
cs.AI updates on arXiv.org