A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
Researchers have introduced the Pair-GRPO family, a new theoretical framework designed to enhance Large Language Model (LLM) alignment via Reinforcement Learning from Human Preferences (RLHF). Addressing common issues like unstable policy updates and high gradient variance in mainstream pairwise preference learning, this framework comprises two variants: Soft-Pair-GRPO and Hard-Pair-GRPO. Soft-Pair-GRPO modifies Group Relative Policy Optimization by using binary pairwise preference rewards, maintaining stability through a proven gradient equivalence theorem. Hard-Pair-GRPO further improves performance by introducing explicit local probability constraints and constrained KL-fitting optimization to suppress gradient noise. Theoretical guarantees include monotonic policy improvement and deterministic gradient direction. Extensive experiments on benchmarks such as HH-RLHF, UltraFeedback, and MuJoCo’s HalfCheetah-v4 demonstrate that the Pair-GRPO family outperforms state-of-the-art baselines in alignment quality, training stability, and generalization. This development marks a significant advancement in making RLHF more robust and interpretable for AI systems.
Wire timeline
A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
Researchers have introduced the Pair-GRPO family, a new theoretical framework designed to enhance Large Language Model (LLM) alignment via Reinforcement Learning from Human Preferences (RLHF). Addressing common issues like unstable policy updates and high gradient variance in mainstream pairwise preference learning, this framework comprises two variants: Soft-Pair-GRPO and Hard-Pair-GRPO. Soft-Pair-GRPO modifies Group Relative Policy Optimization by using binary pairwise preference rewards, maintaining stability through a proven gradient equivalence theorem. Hard-Pair-GRPO further improves performance by introducing explicit local probability constraints and constrained KL-fitting optimization to suppress gradient noise. Theoretical guarantees include monotonic policy improvement and deterministic gradient direction. Extensive experiments on benchmarks such as HH-RLHF, UltraFeedback, and MuJoCo’s HalfCheetah-v4 demonstrate that the Pair-GRPO family outperforms state-of-the-art baselines in alignment quality, training stability, and generalization. This development marks a significant advancement in making RLHF more robust and interpretable for AI systems.
cs.AI updates on arXiv.org