Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
A new research paper published on arXiv reveals that training Large Language Models (LLMs) on mixed datasets can lead to sudden phase transitions in knowledge acquisition, rather than following smooth scaling laws. The study, conducted by researchers including Xinran Gu and Jingzhao Zhang, demonstrates that when models are trained on a mixture of web-scraped data and high-quality, knowledge-dense sources, their ability to memorize specific information changes discontinuously. Experiments using synthetic biography datasets show two critical thresholds: one related to model size and another to the data mixing ratio. Below these critical values, models retain little knowledge; beyond them, retention improves rapidly. The authors attribute this phenomenon to capacity allocation, where bounded models optimize test loss like a knapsack problem solver. An information-theoretic framework formalizes these findings, indicating a power-law relationship between the critical mixing ratio and model size. This research highlights that optimal data mixing strategies differ significantly between small and large models, providing crucial insights for efficient LLM training and resource allocation in artificial intelligence development.
Wire timeline
Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
A new research paper published on arXiv reveals that training Large Language Models (LLMs) on mixed datasets can lead to sudden phase transitions in knowledge acquisition, rather than following smooth scaling laws. The study, conducted by researchers including Xinran Gu and Jingzhao Zhang, demonstrates that when models are trained on a mixture of web-scraped data and high-quality, knowledge-dense sources, their ability to memorize specific information changes discontinuously. Experiments using synthetic biography datasets show two critical thresholds: one related to model size and another to the data mixing ratio. Below these critical values, models retain little knowledge; beyond them, retention improves rapidly. The authors attribute this phenomenon to capacity allocation, where bounded models optimize test loss like a knapsack problem solver. An information-theoretic framework formalizes these findings, indicating a power-law relationship between the critical mixing ratio and model size. This research highlights that optimal data mixing strategies differ significantly between small and large models, providing crucial insights for efficient LLM training and resource allocation in artificial intelligence development.
cs.AI updates on arXiv.org