Spectral Analysis of Layer-wise Gradients Reveals Data Quality Metrics for LLM Post-Training
A new research paper published on arXiv investigates how instruction and reasoning data influence the post-training dynamics of large language models (LLMs). The study introduces a spectral analysis of layer-wise gradients to evaluate data quality, revealing that existing metrics like IFD, InsTag, Difficulty, and Reward can be unified through spectral properties derived from singular value decomposition (SVD). The authors find that high-quality data correlates with lower nuclear norms and higher effective ranks. Notably, effective rank proves more robust than nuclear norm in distinguishing subtle quality differences, with reasoning data exhibiting significantly higher effective ranks than instruction data due to richer gradient structures. Experiments indicate that models within the same family share similar gradient patterns regardless of size, while different families diverge. This work provides a unified framework for understanding the interplay between data quality and training stability, offering novel insights for developing improved data exploration strategies in LLM post-training.
Wire timeline
Spectral Analysis of Layer-wise Gradients Reveals Data Quality Metrics for LLM Post-Training
A new research paper published on arXiv investigates how instruction and reasoning data influence the post-training dynamics of large language models (LLMs). The study introduces a spectral analysis of layer-wise gradients to evaluate data quality, revealing that existing metrics like IFD, InsTag, Difficulty, and Reward can be unified through spectral properties derived from singular value decomposition (SVD). The authors find that high-quality data correlates with lower nuclear norms and higher effective ranks. Notably, effective rank proves more robust than nuclear norm in distinguishing subtle quality differences, with reasoning data exhibiting significantly higher effective ranks than instruction data due to richer gradient structures. Experiments indicate that models within the same family share similar gradient patterns regardless of size, while different families diverge. This work provides a unified framework for understanding the interplay between data quality and training stability, offering novel insights for developing improved data exploration strategies in LLM post-training.
cs.AI updates on arXiv.org