Deciphering Rock Tokens in On-Policy Distillation for Efficient Model Training
A new research paper titled 'Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation' investigates the role of high-loss tokens, termed 'Rock Tokens,' in On-Policy Distillation (OPD). While existing studies suggest such tokens should diminish as training converges, empirical analysis reveals they persist, accounting for up to 18% of generated output tokens even after saturation. The study identifies two paradoxes: Rock Tokens contribute disproportionately to gradient norms yet remain stagnant and resistant to teacher-driven corrections, and they provide negligible functional contribution to the model's reasoning performance through causal intervention. These findings indicate that significant optimization bandwidth is wasted on structural and discourse residuals that student models do not need to internalize. By strategically bypassing these stumbling blocks, the authors demonstrate a method to streamline the alignment process. This challenges the necessity of uniform token weighting and proposes a more efficient paradigm for large-scale model distillation, potentially enhancing the efficiency of training large language models by focusing computational resources on critical tokens rather than persistent, low-value errors.
Wire timeline
Deciphering Rock Tokens in On-Policy Distillation for Efficient Model Training
A new research paper titled 'Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation' investigates the role of high-loss tokens, termed 'Rock Tokens,' in On-Policy Distillation (OPD). While existing studies suggest such tokens should diminish as training converges, empirical analysis reveals they persist, accounting for up to 18% of generated output tokens even after saturation. The study identifies two paradoxes: Rock Tokens contribute disproportionately to gradient norms yet remain stagnant and resistant to teacher-driven corrections, and they provide negligible functional contribution to the model's reasoning performance through causal intervention. These findings indicate that significant optimization bandwidth is wasted on structural and discourse residuals that student models do not need to internalize. By strategically bypassing these stumbling blocks, the authors demonstrate a method to streamline the alignment process. This challenges the necessity of uniform token weighting and proposes a more efficient paradigm for large-scale model distillation, potentially enhancing the efficiency of training large language models by focusing computational resources on critical tokens rather than persistent, low-value errors.
cs.AI updates on arXiv.org