Model Merging Scaling Laws in Large Language Models
Researchers have identified empirical scaling laws for language model merging, addressing the lack of quantitative rules for predicting returns when adding experts or scaling model size. The study reveals a compact power law linking model size and the number of experts, showing that the size-dependent performance floor decreases with capacity while merging exhibits diminishing returns as more experts are added. This law holds across diverse architectures and merging methods, including Average, TA, TIES, and DARE, both in-domain and cross-domain. It explains two key regularities: most performance gains occur early, and variability shrinks with increased expert inclusion. The authors present a theoretical framework explaining why gains decrease roughly as 1/k, connecting these trends to base model properties and domain diversity. This discovery enables predictive planning, allowing developers to estimate the necessary number of experts for target loss levels and optimize budget trade-offs between scaling base models and adding experts. By transforming merging from a heuristic practice into a computationally efficient, planable alternative to multitask training, the findings suggest a scalable principle for distributed generative AI, offering a complementary path toward achieving AGI-level systems through the composition of specialists.
Wire timeline
Model Merging Scaling Laws in Large Language Models
Researchers have identified empirical scaling laws for language model merging, addressing the lack of quantitative rules for predicting returns when adding experts or scaling model size. The study reveals a compact power law linking model size and the number of experts, showing that the size-dependent performance floor decreases with capacity while merging exhibits diminishing returns as more experts are added. This law holds across diverse architectures and merging methods, including Average, TA, TIES, and DARE, both in-domain and cross-domain. It explains two key regularities: most performance gains occur early, and variability shrinks with increased expert inclusion. The authors present a theoretical framework explaining why gains decrease roughly as 1/k, connecting these trends to base model properties and domain diversity. This discovery enables predictive planning, allowing developers to estimate the necessary number of experts for target loss levels and optimize budget trade-offs between scaling base models and adding experts. By transforming merging from a heuristic practice into a computationally efficient, planable alternative to multitask training, the findings suggest a scalable principle for distributed generative AI, offering a complementary path toward achieving AGI-level systems through the composition of specialists.
cs.AI updates on arXiv.org