SDG-MoE: Signed Debate Graph Mixture-of-Experts
Researchers have introduced SDG-MoE, a novel architecture for Sparse Mixture-of-Experts (MoE) models that enhances performance by enabling communication among active experts. Unlike traditional MoE systems where experts process tokens independently, SDG-MoE incorporates a lightweight, iterative deliberation step before final aggregation. This mechanism utilizes two learned interaction matrices—a support graph and a critique graph—to capture reinforcing and corrective influences through signed message-passing. Additionally, it employs disagreement-gated Friedkin-Johnsen-style anchoring to control deliberation strength and prevent expert drift. Theoretical analysis confirms the stability of expert states with minimal computational overhead. In controlled pretraining experiments, SDG-MoE demonstrated significant improvements, outperforming the strongest baseline by 19.8% in validation perplexity. It also achieved superior external perplexity scores on standard benchmarks including WikiText-103, C4, and Paloma. This development addresses the underexplored potential of direct interaction among routed experts, offering a structured approach to preserving specialization while leveraging collective deliberation for improved model accuracy and efficiency in machine learning applications.
Wire timeline
SDG-MoE: Signed Debate Graph Mixture-of-Experts
Researchers have introduced SDG-MoE, a novel architecture for Sparse Mixture-of-Experts (MoE) models that enhances performance by enabling communication among active experts. Unlike traditional MoE systems where experts process tokens independently, SDG-MoE incorporates a lightweight, iterative deliberation step before final aggregation. This mechanism utilizes two learned interaction matrices—a support graph and a critique graph—to capture reinforcing and corrective influences through signed message-passing. Additionally, it employs disagreement-gated Friedkin-Johnsen-style anchoring to control deliberation strength and prevent expert drift. Theoretical analysis confirms the stability of expert states with minimal computational overhead. In controlled pretraining experiments, SDG-MoE demonstrated significant improvements, outperforming the strongest baseline by 19.8% in validation perplexity. It also achieved superior external perplexity scores on standard benchmarks including WikiText-103, C4, and Paloma. This development addresses the underexplored potential of direct interaction among routed experts, offering a structured approach to preserving specialization while leveraging collective deliberation for improved model accuracy and efficiency in machine learning applications.
cs.AI updates on arXiv.org