Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
Researchers have introduced Elastic Mixture-of-Experts (EMoE), a novel training framework designed to enhance the flexibility and efficiency of Mixture-of-Experts (MoE) models during inference. Traditionally, MoE models fix the number of activated experts, limiting their adaptability to diverse hardware and workload requirements. The study identifies an 'inference-time scaling wall,' where increasing activated experts beyond the training limit causes performance degradation due to poor expert collaboration. EMoE addresses this by training experts to collaborate effectively in various combinations and improving router selection quality. Extensive experiments across four MoE architectures (7B–21B parameters) and nine benchmarks demonstrate that EMoE expands the effective scaling range to two or three times the training-time expert count while achieving higher peak performance. This innovation allows a single model to serve diverse quality-latency budgets without the costly need for training separate models for each scenario, significantly advancing the deployment versatility of large-scale AI systems.
Wire timeline
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
Researchers have introduced Elastic Mixture-of-Experts (EMoE), a novel training framework designed to enhance the flexibility and efficiency of Mixture-of-Experts (MoE) models during inference. Traditionally, MoE models fix the number of activated experts, limiting their adaptability to diverse hardware and workload requirements. The study identifies an 'inference-time scaling wall,' where increasing activated experts beyond the training limit causes performance degradation due to poor expert collaboration. EMoE addresses this by training experts to collaborate effectively in various combinations and improving router selection quality. Extensive experiments across four MoE architectures (7B–21B parameters) and nine benchmarks demonstrate that EMoE expands the effective scaling range to two or three times the training-time expert count while achieving higher peak performance. This innovation allows a single model to serve diverse quality-latency budgets without the costly need for training separate models for each scenario, significantly advancing the deployment versatility of large-scale AI systems.
cs.AI updates on arXiv.org