AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning
Researchers have introduced AEM (Adaptive Entropy Modulation), a novel supervision-free credit assignment method designed to enhance reinforcement learning (RL) for large language model (LLM) agents. Addressing the challenge of sparse outcome-only rewards in multi-turn tasks, AEM adaptively modulates entropy dynamics during training to optimize the exploration-exploitation trade-off. Unlike existing approaches that rely on dense intermediate supervision, which increases complexity and limits generalization, AEM operates at the response level rather than the token level. This alignment reduces sensitivity to sampling noise and leverages the interaction between sampled-response advantage and relative surprisal. By deriving a practical response-level uncertainty proxy, the method rescales advantages to facilitate a natural transition from exploration to exploitation. Extensive experiments conducted on benchmarks such as ALFWorld, WebShop, and SWE-bench-Verified, using models ranging from 1.5B to 32B parameters, demonstrate that AEM consistently improves upon strong RL baselines. Notably, it achieved a 1.4% performance gain when integrated into a state-of-the-art software-engineering RL training framework, highlighting its potential for improving agentic capabilities in complex environments.
Wire timeline
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning
Researchers have introduced AEM (Adaptive Entropy Modulation), a novel supervision-free credit assignment method designed to enhance reinforcement learning (RL) for large language model (LLM) agents. Addressing the challenge of sparse outcome-only rewards in multi-turn tasks, AEM adaptively modulates entropy dynamics during training to optimize the exploration-exploitation trade-off. Unlike existing approaches that rely on dense intermediate supervision, which increases complexity and limits generalization, AEM operates at the response level rather than the token level. This alignment reduces sensitivity to sampling noise and leverages the interaction between sampled-response advantage and relative surprisal. By deriving a practical response-level uncertainty proxy, the method rescales advantages to facilitate a natural transition from exploration to exploitation. Extensive experiments conducted on benchmarks such as ALFWorld, WebShop, and SWE-bench-Verified, using models ranging from 1.5B to 32B parameters, demonstrate that AEM consistently improves upon strong RL baselines. Notably, it achieved a 1.4% performance gain when integrated into a state-of-the-art software-engineering RL training framework, highlighting its potential for improving agentic capabilities in complex environments.
cs.AI updates on arXiv.org