Hi-MoE: Hierarchical Mixture-of-Experts with Two-Stage Optimization
Researchers have introduced Hi-MoE, a novel grouped Mixture-of-Experts (MoE) framework designed to address the fundamental trade-off in sparse MoE models between load balancing and expert specialization. Traditional routers often struggle, where strong balancing suppresses specialization, while aggressive diversity leads to routing collapse. Hi-MoE decomposes routing control into two coupled levels: inter-group balancing to ensure fair traffic distribution across expert groups, and intra-group specialization to promote complementary behaviors while preventing within-group collapse. The study provides a theoretical analysis of how these hierarchical objectives stabilize router performance. Empirical results demonstrate consistent improvements over existing sparse-routing and grouped-MoE baselines across natural language processing and computer vision benchmarks. In large-scale pre-training involving 58 billion tokens, the Hi-MoE-7B model achieved a 5.6% reduction in perplexity and a 40% improvement in expert balance compared to the OLMoE-7B baseline. These findings highlight the framework's robustness and efficiency in scaling model capacity without sacrificing performance stability.
Wire timeline
Hi-MoE: Hierarchical Mixture-of-Experts with Two-Stage Optimization
Researchers have introduced Hi-MoE, a novel grouped Mixture-of-Experts (MoE) framework designed to address the fundamental trade-off in sparse MoE models between load balancing and expert specialization. Traditional routers often struggle, where strong balancing suppresses specialization, while aggressive diversity leads to routing collapse. Hi-MoE decomposes routing control into two coupled levels: inter-group balancing to ensure fair traffic distribution across expert groups, and intra-group specialization to promote complementary behaviors while preventing within-group collapse. The study provides a theoretical analysis of how these hierarchical objectives stabilize router performance. Empirical results demonstrate consistent improvements over existing sparse-routing and grouped-MoE baselines across natural language processing and computer vision benchmarks. In large-scale pre-training involving 58 billion tokens, the Hi-MoE-7B model achieved a 5.6% reduction in perplexity and a 40% improvement in expert balance compared to the OLMoE-7B baseline. These findings highlight the framework's robustness and efficiency in scaling model capacity without sacrificing performance stability.
cs.AI updates on arXiv.org