PrAg-PO: Prompt Augmented Policy Optimization for Robust Mathematical Reasoning
Researchers have introduced Prompt Augmented Policy Optimization (PrAg-PO), a new reinforcement learning method designed to enhance the mathematical reasoning capabilities of large language models. While existing algorithms like Group-Relative Policy Optimization (GRPO) often suffer from prompt-specific overfitting due to reliance on single fixed templates, PrAg-PO addresses this by mixing diverse prompt templates with template-specific format rewards during training. This approach significantly increases rollout diversity and improves model robustness, effectively mitigating premature training collapse. Empirical evaluations conducted on models such as DeepSeek-R1-Distill-Qwen-1.5B, Qwen2.5-Math-1.5B, and Qwen3-1.7B demonstrate that PrAg-PO consistently outperforms strong baselines, including GRPO and DAPO. The method achieves higher reasoning accuracy on mathematics benchmarks using only a fixed training set of 8.5K problems from MATH Level 3-5. The study highlights PrAg-PO as a simple yet effective solution for stabilizing training dynamics and boosting performance in complex reasoning tasks. Code and model checkpoints for this innovation are publicly available, marking a significant step forward in optimizing policy methods for artificial intelligence applications in mathematics.
Wire timeline
PrAg-PO: Prompt Augmented Policy Optimization for Robust Mathematical Reasoning
Researchers have introduced Prompt Augmented Policy Optimization (PrAg-PO), a new reinforcement learning method designed to enhance the mathematical reasoning capabilities of large language models. While existing algorithms like Group-Relative Policy Optimization (GRPO) often suffer from prompt-specific overfitting due to reliance on single fixed templates, PrAg-PO addresses this by mixing diverse prompt templates with template-specific format rewards during training. This approach significantly increases rollout diversity and improves model robustness, effectively mitigating premature training collapse. Empirical evaluations conducted on models such as DeepSeek-R1-Distill-Qwen-1.5B, Qwen2.5-Math-1.5B, and Qwen3-1.7B demonstrate that PrAg-PO consistently outperforms strong baselines, including GRPO and DAPO. The method achieves higher reasoning accuracy on mathematics benchmarks using only a fixed training set of 8.5K problems from MATH Level 3-5. The study highlights PrAg-PO as a simple yet effective solution for stabilizing training dynamics and boosting performance in complex reasoning tasks. Code and model checkpoints for this innovation are publicly available, marking a significant step forward in optimizing policy methods for artificial intelligence applications in mathematics.
cs.AI updates on arXiv.org