HTPO: Hierarchical Token-level Objective Control for Balanced RL in LLMs
Researchers have introduced Hierarchical Token-level Objective Control Policy Optimization (HTPO), a novel reinforcement learning algorithm designed to enhance the reasoning capabilities of Large Language Models (LLMs). Current methods often treat all response tokens equally, failing to balance exploration and exploitation effectively during Chain-of-Thought reasoning. HTPO addresses this by hierarchically partitioning tokens into functional groups based on prompt difficulty, answer correctness, and token entropy. It applies specialized optimization objectives to each group, allowing for granular control over the learning process. Extensive experiments on challenging reasoning benchmarks demonstrate that HTPO significantly outperforms the strong DAPO baseline, achieving improvements of 8.6% on AIME'24 and 6.7% on AIME'25. Furthermore, the model maintains a consistent performance advantage as test-time compute scales, indicating that adaptive token-level control fosters effective exploration without compromising exploitation. This advancement highlights a significant step forward in optimizing LLM reasoning through more sophisticated reinforcement learning techniques.
Wire timeline
HTPO: Hierarchical Token-level Objective Control for Balanced RL in LLMs
Researchers have introduced Hierarchical Token-level Objective Control Policy Optimization (HTPO), a novel reinforcement learning algorithm designed to enhance the reasoning capabilities of Large Language Models (LLMs). Current methods often treat all response tokens equally, failing to balance exploration and exploitation effectively during Chain-of-Thought reasoning. HTPO addresses this by hierarchically partitioning tokens into functional groups based on prompt difficulty, answer correctness, and token entropy. It applies specialized optimization objectives to each group, allowing for granular control over the learning process. Extensive experiments on challenging reasoning benchmarks demonstrate that HTPO significantly outperforms the strong DAPO baseline, achieving improvements of 8.6% on AIME'24 and 6.7% on AIME'25. Furthermore, the model maintains a consistent performance advantage as test-time compute scales, indicating that adaptive token-level control fosters effective exploration without compromising exploitation. This advancement highlights a significant step forward in optimizing LLM reasoning through more sophisticated reinforcement learning techniques.
cs.AI updates on arXiv.org