Continual Harness: Online Adaptation for Self-Improving Foundation Agents
Researchers have introduced Continual Harness, a novel framework enabling embodied AI agents to self-improve through online adaptation without requiring environment resets. Unlike existing coding harnesses that rely on human-in-the-loop refinement or episode resets for prompt optimization, this system allows agents to iteratively refine their prompts, sub-agents, skills, and memory during a single continuous run. The study highlights initial experiments with Gemini Plays Pokemon (GPP), which achieved perfect runs in Pokemon Blue, Yellow Legacy, and Crystal. Continual Harness automates this process, demonstrating significant reductions in button-press costs and recovering much of the performance gap compared to hand-engineered expert systems in Pokemon Red and Emerald. Furthermore, the researchers implemented an online process-reward co-learning loop where a frontier teacher model relabels rollouts to update open-source agents. This approach drives sustained progress in complex tasks despite starting with minimal interface knowledge and no curated domain scaffolding, marking a significant advancement in long-horizon decision-making for foundation models.
Wire timeline
Continual Harness: Online Adaptation for Self-Improving Foundation Agents
Researchers have introduced Continual Harness, a novel framework enabling embodied AI agents to self-improve through online adaptation without requiring environment resets. Unlike existing coding harnesses that rely on human-in-the-loop refinement or episode resets for prompt optimization, this system allows agents to iteratively refine their prompts, sub-agents, skills, and memory during a single continuous run. The study highlights initial experiments with Gemini Plays Pokemon (GPP), which achieved perfect runs in Pokemon Blue, Yellow Legacy, and Crystal. Continual Harness automates this process, demonstrating significant reductions in button-press costs and recovering much of the performance gap compared to hand-engineered expert systems in Pokemon Red and Emerald. Furthermore, the researchers implemented an online process-reward co-learning loop where a frontier teacher model relabels rollouts to update open-source agents. This approach drives sustained progress in complex tasks despite starting with minimal interface knowledge and no curated domain scaffolding, marking a significant advancement in long-horizon decision-making for foundation models.
cs.AI updates on arXiv.org