ASTOR: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs
Researchers have introduced ASTOR, a novel multi-task reinforcement learning framework designed to enhance the post-training of Large Language Models (LLMs) for coding tasks. While reinforcement learning with verifiable rewards is effective, existing multi-task approaches often treat all coding tasks uniformly, limiting training efficiency. ASTOR addresses this by utilizing task utility signals to capture learning potential and cross-task synergy. The framework features two core modules: a Hierarchical Utility-Routed Data Scheduling module that prioritizes informative prompts and allocates training budgets, and an Adaptive Utility-Calibrated Policy Optimization module that dynamically adjusts KL regularization based on each task's current state. Experimental results across four representative coding tasks using two widely-used LLMs demonstrate that ASTOR significantly outperforms both task-specific specialists and strong multi-task baselines. Specifically, it achieves a 9.0%-9.5% improvement over the best single-task models and surpasses the strongest multi-task baseline by 7.5%-12.8%, offering a more cost-effective and powerful solution for unified code generation models.
Wire timeline
ASTOR: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs
Researchers have introduced ASTOR, a novel multi-task reinforcement learning framework designed to enhance the post-training of Large Language Models (LLMs) for coding tasks. While reinforcement learning with verifiable rewards is effective, existing multi-task approaches often treat all coding tasks uniformly, limiting training efficiency. ASTOR addresses this by utilizing task utility signals to capture learning potential and cross-task synergy. The framework features two core modules: a Hierarchical Utility-Routed Data Scheduling module that prioritizes informative prompts and allocates training budgets, and an Adaptive Utility-Calibrated Policy Optimization module that dynamically adjusts KL regularization based on each task's current state. Experimental results across four representative coding tasks using two widely-used LLMs demonstrate that ASTOR significantly outperforms both task-specific specialists and strong multi-task baselines. Specifically, it achieves a 9.0%-9.5% improvement over the best single-task models and surpasses the strongest multi-task baseline by 7.5%-12.8%, offering a more cost-effective and powerful solution for unified code generation models.
cs.AI updates on arXiv.org