Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
Researchers have introduced Asymmetric On-Policy Distillation (AOPD), a novel machine learning technique designed to enhance the training of student models using token-level teacher feedback. While standard On-Policy Distillation (OPD) often outperforms other methods, it suffers from structural weaknesses such as high variance updates, vanishing gradients in zero-advantage regions, and exploration bottlenecks. AOPD addresses these issues by replacing ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions, while preserving positive reinforcement learning mechanisms. Experimental results on mathematical reasoning benchmarks demonstrate that AOPD consistently surpasses standard OPD, achieving average performance gains of 4.09 under strong initialization and 8.34 under weak initialization. Furthermore, the method maintains higher policy entropy during training and ensures better capability retention during sequential tool-use adaptation. This development represents a significant advancement in optimizing reinforcement learning strategies for complex reasoning tasks, offering a more robust framework for balancing exploitation and imitation in artificial intelligence systems.
Wire timeline
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
Researchers have introduced Asymmetric On-Policy Distillation (AOPD), a novel machine learning technique designed to enhance the training of student models using token-level teacher feedback. While standard On-Policy Distillation (OPD) often outperforms other methods, it suffers from structural weaknesses such as high variance updates, vanishing gradients in zero-advantage regions, and exploration bottlenecks. AOPD addresses these issues by replacing ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions, while preserving positive reinforcement learning mechanisms. Experimental results on mathematical reasoning benchmarks demonstrate that AOPD consistently surpasses standard OPD, achieving average performance gains of 4.09 under strong initialization and 8.34 under weak initialization. Furthermore, the method maintains higher policy entropy during training and ensures better capability retention during sequential tool-use adaptation. This development represents a significant advancement in optimizing reinforcement learning strategies for complex reasoning tasks, offering a more robust framework for balancing exploitation and imitation in artificial intelligence systems.
cs.AI updates on arXiv.org