Accelerating LMO-Based Optimization via Implicit Gradient Transport
Researchers Won-Jun Jang and Si-Hyeon Lee have proposed LMO-IGT, a new class of stochastic optimization methods leveraging implicit gradient transport (IGT) to accelerate Linear Minimization Oracle (LMO)-based optimizers like Lion and Muon. While existing variance reduction techniques improve convergence, they often require additional gradient evaluations, increasing computational overhead. LMO-IGT addresses this by evaluating stochastic gradients at transported points, achieving an iteration complexity of O(ε^-3.5) while maintaining the efficient single-gradient-per-iteration structure. This performance bridges the gap between standard stochastic LMO (O(ε^-4)) and variance-reduced methods (O(ε^-3)). The study also introduces a unified theoretical framework and a new stationarity measure called the regularized support function (RSF), which connects gradient-norm and Frank-Wolfe-gap concepts. Empirical results demonstrate that LMO-IGT consistently outperforms standard stochastic LMO counterparts with negligible overhead. Specifically, the Muon-IGT instantiation showed the strongest overall performance across evaluated settings, highlighting IGT as a practical and effective mechanism for enhancing modern LMO-based optimization in machine learning applications.
Wire timeline
Accelerating LMO-Based Optimization via Implicit Gradient Transport
Researchers Won-Jun Jang and Si-Hyeon Lee have proposed LMO-IGT, a new class of stochastic optimization methods leveraging implicit gradient transport (IGT) to accelerate Linear Minimization Oracle (LMO)-based optimizers like Lion and Muon. While existing variance reduction techniques improve convergence, they often require additional gradient evaluations, increasing computational overhead. LMO-IGT addresses this by evaluating stochastic gradients at transported points, achieving an iteration complexity of O(ε^-3.5) while maintaining the efficient single-gradient-per-iteration structure. This performance bridges the gap between standard stochastic LMO (O(ε^-4)) and variance-reduced methods (O(ε^-3)). The study also introduces a unified theoretical framework and a new stationarity measure called the regularized support function (RSF), which connects gradient-norm and Frank-Wolfe-gap concepts. Empirical results demonstrate that LMO-IGT consistently outperforms standard stochastic LMO counterparts with negligible overhead. Specifically, the Muon-IGT instantiation showed the strongest overall performance across evaluated settings, highlighting IGT as a practical and effective mechanism for enhancing modern LMO-based optimization in machine learning applications.
cs.AI updates on arXiv.org