DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
Researchers have introduced DUET (Dual-controlled tokEn allocaTion), a novel method designed to optimize token-budget allocation in Reinforcement Learning with Verifiable Rewards (RLVR). RLVR training is computationally expensive, primarily due rollout generation. Unlike prior works that control either prompt selection or rollout length independently, DUET jointly tunes both dimensions under a shared compute budget. It employs a lightweight pre-rollout surrogate to determine rollout quantity per prompt and a marker-gated abort rule with importance reweighting to decide rollout termination. Experimental results on Qwen3-1.7B trained on the MATH benchmark demonstrate that DUET outperforms full-budget GRPO and other baseline methods in reasoning quality while achieving a 1.62x wall-clock speedup. Notably, using only 50% of the token budget, DUET surpasses all baselines operating at full budget, delivering a 2.51x speedup. The method's effectiveness generalizes across math, coding, and scientific Q&A domains and is verified on other models like Qwen3-4B and Llama-3.2-3B-Instruct. Crucially, DUET's performance advantage widens as budgets tighten, challenging the typical trade-off between efficiency and quality, suggesting it enhances both training speed and learning signal quality.
Wire timeline
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
Researchers have introduced DUET (Dual-controlled tokEn allocaTion), a novel method designed to optimize token-budget allocation in Reinforcement Learning with Verifiable Rewards (RLVR). RLVR training is computationally expensive, primarily due rollout generation. Unlike prior works that control either prompt selection or rollout length independently, DUET jointly tunes both dimensions under a shared compute budget. It employs a lightweight pre-rollout surrogate to determine rollout quantity per prompt and a marker-gated abort rule with importance reweighting to decide rollout termination. Experimental results on Qwen3-1.7B trained on the MATH benchmark demonstrate that DUET outperforms full-budget GRPO and other baseline methods in reasoning quality while achieving a 1.62x wall-clock speedup. Notably, using only 50% of the token budget, DUET surpasses all baselines operating at full budget, delivering a 2.51x speedup. The method's effectiveness generalizes across math, coding, and scientific Q&A domains and is verified on other models like Qwen3-4B and Llama-3.2-3B-Instruct. Crucially, DUET's performance advantage widens as budgets tighten, challenging the typical trade-off between efficiency and quality, suggesting it enhances both training speed and learning signal quality.
cs.AI updates on arXiv.org