Workspace Optimization: A New Method for Training AI Agents Without Weight Updates
Researchers have introduced a novel concept called 'workspace optimization' to enhance the performance of AI agents built on frontier language models that cannot adapt their internal weights. Published on arXiv, the paper argues that since model weights are static, the agent's external workspace—where it reads, writes, and tests information—should be the target of training. This approach mirrors traditional weight-space training by treating artifacts as parameters, evidence as data, counterexamples as losses, and textual feedback as gradients. The authors instantiated this theory in 'DreamTeam,' a multi-agent system designed for the ARC-AGI-3 benchmark. DreamTeam employs specialized roles to build world models, plan, hypothesize, and strategize. In tests on the public 25-game ARC-AGI-3 set, DreamTeam improved the state-of-the-art score from 36% to 38.4% while reducing environment actions by 31%. This development highlights a shift towards optimizing external interaction structures rather than internal model parameters for complex, multi-turn tasks.
Wire timeline
Workspace Optimization: A New Method for Training AI Agents Without Weight Updates
Researchers have introduced a novel concept called 'workspace optimization' to enhance the performance of AI agents built on frontier language models that cannot adapt their internal weights. Published on arXiv, the paper argues that since model weights are static, the agent's external workspace—where it reads, writes, and tests information—should be the target of training. This approach mirrors traditional weight-space training by treating artifacts as parameters, evidence as data, counterexamples as losses, and textual feedback as gradients. The authors instantiated this theory in 'DreamTeam,' a multi-agent system designed for the ARC-AGI-3 benchmark. DreamTeam employs specialized roles to build world models, plan, hypothesize, and strategize. In tests on the public 25-game ARC-AGI-3 set, DreamTeam improved the state-of-the-art score from 36% to 38.4% while reducing environment actions by 31%. This development highlights a shift towards optimizing external interaction structures rather than internal model parameters for complex, multi-turn tasks.
cs.AI updates on arXiv.org