Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
Researchers have introduced Evolving-RL, a novel algorithmic framework designed to enhance the self-evolving capabilities of large language model (LLM) agents. Addressing the static nature of current models, this approach focuses on jointly optimizing experience extraction and utilization through reinforcement learning. Unlike previous studies that primarily addressed system-level design or isolated optimization stages, Evolving-RL treats self-evolution as a unified process. It employs supervisory signals from evaluation to coordinate the co-evolution of extractors and solvers. Experimental results on ALFWorld and Mind2Web benchmarks demonstrate significant performance improvements, including a 98.7% relative gain over the GRPO baseline on unseen ALFWorld tasks and a 35.8% improvement on Mind2Web. The framework effectively internalizes reusable experience patterns into model parameters, enabling robust adaptation to out-of-distribution tasks even without test-time experience accumulation. This advancement highlights a critical shift towards optimizing the inherent abstraction and generalization capacities of foundation models for dynamic deployment scenarios.
Wire timeline
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
Researchers have introduced Evolving-RL, a novel algorithmic framework designed to enhance the self-evolving capabilities of large language model (LLM) agents. Addressing the static nature of current models, this approach focuses on jointly optimizing experience extraction and utilization through reinforcement learning. Unlike previous studies that primarily addressed system-level design or isolated optimization stages, Evolving-RL treats self-evolution as a unified process. It employs supervisory signals from evaluation to coordinate the co-evolution of extractors and solvers. Experimental results on ALFWorld and Mind2Web benchmarks demonstrate significant performance improvements, including a 98.7% relative gain over the GRPO baseline on unseen ALFWorld tasks and a 35.8% improvement on Mind2Web. The framework effectively internalizes reusable experience patterns into model parameters, enabling robust adaptation to out-of-distribution tasks even without test-time experience accumulation. This advancement highlights a critical shift towards optimizing the inherent abstraction and generalization capacities of foundation models for dynamic deployment scenarios.
cs.AI updates on arXiv.org