BEACON Framework Enhances Long-Horizon Language Agent Training
Researchers have introduced BEACON, a novel milestone-guided policy learning framework designed to address significant challenges in training language agents for long-horizon tasks using reinforcement learning. The study identifies credit misattribution and sample inefficiency as primary obstacles, where correct early actions are often penalized by terminal failures, and scarce successful trajectories lead to lost learning signals. BEACON leverages the compositional structure of complex tasks by partitioning trajectories at milestone boundaries, applying temporal reward shaping to credit partial progress, and estimating advantages at dual scales. This approach prevents distant failures from corrupting local action evaluations. Experimental results on benchmarks such as ALFWorld, WebShop, and ScienceWorld demonstrate that BEACON consistently outperforms existing methods like GRPO and GiGPO. Notably, on long-horizon ALFWorld tasks, BEACON achieved a 92.9% success rate, nearly doubling GRPO's performance, while significantly improving effective sample utilization from 23.7% to 82.0%. These findings establish milestone-anchored credit assignment as an effective paradigm for advancing the capabilities of autonomous language agents in complex, sequential decision-making environments.
Wire timeline
BEACON Framework Enhances Long-Horizon Language Agent Training
Researchers have introduced BEACON, a novel milestone-guided policy learning framework designed to address significant challenges in training language agents for long-horizon tasks using reinforcement learning. The study identifies credit misattribution and sample inefficiency as primary obstacles, where correct early actions are often penalized by terminal failures, and scarce successful trajectories lead to lost learning signals. BEACON leverages the compositional structure of complex tasks by partitioning trajectories at milestone boundaries, applying temporal reward shaping to credit partial progress, and estimating advantages at dual scales. This approach prevents distant failures from corrupting local action evaluations. Experimental results on benchmarks such as ALFWorld, WebShop, and ScienceWorld demonstrate that BEACON consistently outperforms existing methods like GRPO and GiGPO. Notably, on long-horizon ALFWorld tasks, BEACON achieved a 92.9% success rate, nearly doubling GRPO's performance, while significantly improving effective sample utilization from 23.7% to 82.0%. These findings establish milestone-anchored credit assignment as an effective paradigm for advancing the capabilities of autonomous language agents in complex, sequential decision-making environments.
cs.AI updates on arXiv.org