Optimizer-Induced Mode Connectivity: From AdamW to Muon
A new research paper submitted to arXiv investigates the underexplored role of optimizers in mode connectivity within neural networks. The study, titled 'Optimizer-Induced Mode Connectivity: From AdamW to Muon,' examines how implicit regularization from specific optimizers affects the connectivity of solution sets. For two-layer ReLU networks, the authors demonstrate that solutions derived from a single optimizer, such as AdamW or Muon, form a connected set at sufficiently large widths. This finding extends beyond prior literature by highlighting optimizer-dependent structures. The research further characterizes interactions between different optimizer-induced regions, noting that they may be disjoint or overlapping depending on regularization strength. In small-width scenarios, AdamW and Muon converge to disconnected zero-loss components separated by a provable loss barrier. Empirical evidence from GPT-2 pretraining supports these theoretical insights, showing that same-optimizer paths preserve model spectra, whereas cross-optimizer paths facilitate smooth transitions. These results provide deeper understanding of optimization landscapes in deep learning.
Wire timeline
Optimizer-Induced Mode Connectivity: From AdamW to Muon
A new research paper submitted to arXiv investigates the underexplored role of optimizers in mode connectivity within neural networks. The study, titled 'Optimizer-Induced Mode Connectivity: From AdamW to Muon,' examines how implicit regularization from specific optimizers affects the connectivity of solution sets. For two-layer ReLU networks, the authors demonstrate that solutions derived from a single optimizer, such as AdamW or Muon, form a connected set at sufficiently large widths. This finding extends beyond prior literature by highlighting optimizer-dependent structures. The research further characterizes interactions between different optimizer-induced regions, noting that they may be disjoint or overlapping depending on regularization strength. In small-width scenarios, AdamW and Muon converge to disconnected zero-loss components separated by a provable loss barrier. Empirical evidence from GPT-2 pretraining supports these theoretical insights, showing that same-optimizer paths preserve model spectra, whereas cross-optimizer paths facilitate smooth transitions. These results provide deeper understanding of optimization landscapes in deep learning.
cs.AI updates on arXiv.org