Improving Generalization by Permutation Routing Across Model Copies
Researchers Shuhei Kashiwamura and Timothee Leleu have introduced a novel machine learning technique utilizing the M-cover transform to enhance model generalization. Published on arXiv in May 2026, this method involves replicating a model M times but diverges from traditional approaches like replicated SGD or Elastic SGD. Instead of coupling copies through parameter averaging or explicit attractive forces, the new framework rewires the contexts for computing local learning messages. Local losses are evaluated on routed models with parameters drawn from different copies via permutations sampled from a structured mixing kernel Q. This kernel defines the topology for message transport and controls the long-loop structure of the lifted factor graph. The study demonstrates that this principle applies effectively across various architectures, including perceptrons, committee machines, and multilayer perceptrons. By facilitating structured message sharing rather than relying on replica collapse or parameter-space coupling, the proposed framework offers a robust mechanism for improving generalization in both discrete models and differentiable neural networks, marking a significant advancement in artificial intelligence training methodologies.
Wire timeline
Improving Generalization by Permutation Routing Across Model Copies
Researchers Shuhei Kashiwamura and Timothee Leleu have introduced a novel machine learning technique utilizing the M-cover transform to enhance model generalization. Published on arXiv in May 2026, this method involves replicating a model M times but diverges from traditional approaches like replicated SGD or Elastic SGD. Instead of coupling copies through parameter averaging or explicit attractive forces, the new framework rewires the contexts for computing local learning messages. Local losses are evaluated on routed models with parameters drawn from different copies via permutations sampled from a structured mixing kernel Q. This kernel defines the topology for message transport and controls the long-loop structure of the lifted factor graph. The study demonstrates that this principle applies effectively across various architectures, including perceptrons, committee machines, and multilayer perceptrons. By facilitating structured message sharing rather than relying on replica collapse or parameter-space coupling, the proposed framework offers a robust mechanism for improving generalization in both discrete models and differentiable neural networks, marking a significant advancement in artificial intelligence training methodologies.
cs.AI updates on arXiv.org