LEPO: Latent Reasoning Policy Optimization for Large Language Models
Researchers have introduced LEPO (Latent Reasoning Policy Optimization), a novel framework designed to enhance the reasoning capabilities of Large Language Models (LLMs) through Reinforcement Learning. While recent advancements in latent reasoning allow LLMs to utilize continuous space information, these methods often suffer from deterministic inference collapse, limiting the discovery of diverse reasoning paths. To address this, the authors inject controllable stochasticity into the latent reasoning process using Gumbel-Softmax techniques. This approach restores the model's exploratory capacity and improves compatibility with Reinforcement Learning algorithms. The LEPO framework applies RL directly to continuous latent representations, maintaining stochasticity during the rollout stage for diverse trajectory sampling. Furthermore, it constructs a unified gradient estimation for both latent representations and discrete tokens during the optimization stage. Extensive experimental results demonstrate that LEPO significantly outperforms existing Reinforcement Learning methods used for both discrete and latent reasoning tasks, marking a significant advancement in optimizing LLM performance through continuous space exploration.
Wire timeline
LEPO: Latent Reasoning Policy Optimization for Large Language Models
Researchers have introduced LEPO (Latent Reasoning Policy Optimization), a novel framework designed to enhance the reasoning capabilities of Large Language Models (LLMs) through Reinforcement Learning. While recent advancements in latent reasoning allow LLMs to utilize continuous space information, these methods often suffer from deterministic inference collapse, limiting the discovery of diverse reasoning paths. To address this, the authors inject controllable stochasticity into the latent reasoning process using Gumbel-Softmax techniques. This approach restores the model's exploratory capacity and improves compatibility with Reinforcement Learning algorithms. The LEPO framework applies RL directly to continuous latent representations, maintaining stochasticity during the rollout stage for diverse trajectory sampling. Furthermore, it constructs a unified gradient estimation for both latent representations and discrete tokens during the optimization stage. Extensive experimental results demonstrate that LEPO significantly outperforms existing Reinforcement Learning methods used for both discrete and latent reasoning tasks, marking a significant advancement in optimizing LLM performance through continuous space exploration.
cs.AI updates on arXiv.org