EXPO: Exploration-Prioritized Policy Optimization for LLM Mathematical Reasoning
Researchers have introduced EXPO (Exploration-Prioritized Policy Optimization), a novel algorithm designed to enhance Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs) in mathematical reasoning. Addressing inefficiencies in the mainstream Group Relative Policy Optimization (GRPO) method, EXPO targets two key issues: restrictive fixed KL penalty coefficients and uniform question sampling that ignores problem difficulty. The proposed solution incorporates two lightweight modules: Accuracy-Conditioned KL Scaling (AKL), which dynamically adjusts regularization strength based on batch accuracy, and Gaussian Curriculum Sampling (GCS), which prioritizes moderately difficult problems to optimize gradient signals. Extensive experiments conducted on DeepSeek-R1-Distill-Qwen-1.5B and Qwen3-8B-Base models across six benchmarks demonstrate that EXPO consistently outperforms vanilla GRPO. Notably, it achieved a significant absolute gain of 13.34% on the AIME 2025 pass@32 metric, improving from 63.33% to 76.67%. These results indicate that EXPO effectively expands the model's exploration boundary under fixed inference costs, offering a substantial advancement in optimizing LLM performance for complex logical tasks.
Wire timeline
EXPO: Exploration-Prioritized Policy Optimization for LLM Mathematical Reasoning
Researchers have introduced EXPO (Exploration-Prioritized Policy Optimization), a novel algorithm designed to enhance Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs) in mathematical reasoning. Addressing inefficiencies in the mainstream Group Relative Policy Optimization (GRPO) method, EXPO targets two key issues: restrictive fixed KL penalty coefficients and uniform question sampling that ignores problem difficulty. The proposed solution incorporates two lightweight modules: Accuracy-Conditioned KL Scaling (AKL), which dynamically adjusts regularization strength based on batch accuracy, and Gaussian Curriculum Sampling (GCS), which prioritizes moderately difficult problems to optimize gradient signals. Extensive experiments conducted on DeepSeek-R1-Distill-Qwen-1.5B and Qwen3-8B-Base models across six benchmarks demonstrate that EXPO consistently outperforms vanilla GRPO. Notably, it achieved a significant absolute gain of 13.34% on the AIME 2025 pass@32 metric, improving from 63.33% to 76.67%. These results indicate that EXPO effectively expands the model's exploration boundary under fixed inference costs, offering a substantial advancement in optimizing LLM performance for complex logical tasks.
cs.AI updates on arXiv.org