Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
Researchers have introduced a novel training-free diagnostic framework to analyze on-policy distillation in reasoning models, addressing uncertainties about when this technique is beneficial or detrimental. The study defines an ideal per-node gradient that maximally increases a student model's success probability and employs a scalable targeted-rollout algorithm to estimate it efficiently. By calculating a gradient alignment score, the team quantifies how closely distillation signals match this ideal. Key findings reveal that distillation guidance aligns better with ideal gradients during incorrect rollouts, whereas it tends to become noisy when students already perform well. Furthermore, the optimal distillation context varies based on the student model's capacity and the specific task, indicating no universal configuration exists. This research highlights the necessity for per-task and per-token diagnostic analyses to optimize distillation strategies, moving beyond aggregate performance metrics that obscure individual token dynamics. The work provides critical insights for improving the efficiency and effectiveness of training large language models through self-distillation and external teacher models.
Wire timeline
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
Researchers have introduced a novel training-free diagnostic framework to analyze on-policy distillation in reasoning models, addressing uncertainties about when this technique is beneficial or detrimental. The study defines an ideal per-node gradient that maximally increases a student model's success probability and employs a scalable targeted-rollout algorithm to estimate it efficiently. By calculating a gradient alignment score, the team quantifies how closely distillation signals match this ideal. Key findings reveal that distillation guidance aligns better with ideal gradients during incorrect rollouts, whereas it tends to become noisy when students already perform well. Furthermore, the optimal distillation context varies based on the student model's capacity and the specific task, indicating no universal configuration exists. This research highlights the necessity for per-task and per-token diagnostic analyses to optimize distillation strategies, moving beyond aggregate performance metrics that obscure individual token dynamics. The work provides critical insights for improving the efficiency and effectiveness of training large language models through self-distillation and external teacher models.
cs.AI updates on arXiv.org