SlimQwen: Exploring Pruning and Distillation in Large MoE Model Pre-training
Researchers have released a new study titled 'SlimQwen,' which systematically investigates the compression of large Mixture-of-Experts (MoE) language models during pre-training. The paper addresses key uncertainties regarding structured pruning and knowledge distillation (KD) at scale. Key findings indicate that pruning a pretrained MoE model yields better initialization than training from scratch under identical budgets. The study introduces a partial-preservation expert merging strategy that enhances downstream performance across various benchmarks. Additionally, combining KD with language modeling loss proves more effective than KD alone, particularly for knowledge-intensive tasks, with multi-token prediction distillation offering consistent improvements. The research demonstrates that progressive pruning schedules outperform one-shot compression by facilitating better optimization trajectories. As a practical application, the team successfully compressed the Qwen3-Next-80A3B model into a significantly smaller 23A2B model while retaining competitive performance. These results provide valuable guidance for efficient large-scale MoE compression, contributing to the development of more resource-efficient artificial intelligence systems.
Wire timeline
SlimQwen: Exploring Pruning and Distillation in Large MoE Model Pre-training
Researchers have released a new study titled 'SlimQwen,' which systematically investigates the compression of large Mixture-of-Experts (MoE) language models during pre-training. The paper addresses key uncertainties regarding structured pruning and knowledge distillation (KD) at scale. Key findings indicate that pruning a pretrained MoE model yields better initialization than training from scratch under identical budgets. The study introduces a partial-preservation expert merging strategy that enhances downstream performance across various benchmarks. Additionally, combining KD with language modeling loss proves more effective than KD alone, particularly for knowledge-intensive tasks, with multi-token prediction distillation offering consistent improvements. The research demonstrates that progressive pruning schedules outperform one-shot compression by facilitating better optimization trajectories. As a practical application, the team successfully compressed the Qwen3-Next-80A3B model into a significantly smaller 23A2B model while retaining competitive performance. These results provide valuable guidance for efficient large-scale MoE compression, contributing to the development of more resource-efficient artificial intelligence systems.
cs.AI updates on arXiv.org