Kwai AI's SRPO Framework Achieves 10x Efficiency Over GRPO in LLM Training
Researchers from Kwai AI have introduced SRPO (Two-Staged history-Resampling Policy Optimization), a novel reinforcement learning framework that significantly enhances the efficiency of training large language models. Addressing limitations in standard GRPO methods, such as cross-domain optimization conflicts and inefficient sample utilization, SRPO employs a two-stage approach. The first stage focuses on mathematical data to elicit complex reasoning abilities, while the second integrates code data for skill consolidation. This method allows the model to achieve performance levels comparable to DeepSeek-R1-Zero in both math and coding benchmarks, specifically AIME24 and LiveCodeBench, using only one-tenth of the training steps. The team has open-sourced the SRPO-Qwen-32B model and released a technical report detailing their findings. This breakthrough suggests that specialized RL strategies can overcome performance bottlenecks and premature saturation often seen in mixed-domain datasets, marking a significant advancement in optimizing post-training processes for sophisticated reasoning behaviors in LLMs without requiring massive computational resources.
Wire timeline
Kwai AI's SRPO Framework Achieves 10x Efficiency Over GRPO in LLM Training
Researchers from Kwai AI have introduced SRPO (Two-Staged history-Resampling Policy Optimization), a novel reinforcement learning framework that significantly enhances the efficiency of training large language models. Addressing limitations in standard GRPO methods, such as cross-domain optimization conflicts and inefficient sample utilization, SRPO employs a two-stage approach. The first stage focuses on mathematical data to elicit complex reasoning abilities, while the second integrates code data for skill consolidation. This method allows the model to achieve performance levels comparable to DeepSeek-R1-Zero in both math and coding benchmarks, specifically AIME24 and LiveCodeBench, using only one-tenth of the training steps. The team has open-sourced the SRPO-Qwen-32B model and released a technical report detailing their findings. This breakthrough suggests that specialized RL strategies can overcome performance bottlenecks and premature saturation often seen in mixed-domain datasets, marking a significant advancement in optimizing post-training processes for sophisticated reasoning behaviors in LLMs without requiring massive computational resources.
Synced