vix.ing · top · new · best · stats · spec

Muon is Provably Faster with Momentum Variance Reduction

2025/12/18 by Qian, Xun, Rammal, Hussein, Kovalev, Dmitry +1 · 1 citation
Computer Science · Physics and Astronomy · #Computational Physics and Python Applications #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Optimization and Control (math.OC) #Particle physics theoretical and experimental studies #Stochastic Gradient Optimization Techniques

paper · pdf · doi:10.48550/arxiv.2512.16598

openalex publication_date 2025/12/18 · openalex created_date 2025/12/21 · openalex updated_date 2026/07/28

Abstract

Recent empirical research has demonstrated that deep learning optimizers based on the linear minimization oracle (LMO) over specifically chosen Non-Euclidean norm balls, such as Muon and Scion, outperform Adam-type methods in the training of large language models. In this work, we show that such optimizers can be provably improved by replacing their vanilla momentum by momentum variance reduction (MVR). Instead of proposing and analyzing MVR variants of Muon and Scion separately, we incorporate MVR into the recently proposed Gluon framework, which captures Muon, Scion and other specific Non-Euclidean LMO-based methods as special cases, and at the same time works with a more general smoothness assumption which better captures the layer-wise structure of neural networks. In the non-convex case, we incorporate MVR into Gluon in three different ways. All of them improve the convergence rate from \cal O (\frac1K1/4) to \cal O (\frac1K1/3). Additionally, we provide improved rates in the star-convex case. Finally, we conduct several numerical experiments that verify the superior performance of our proposed algorithms in terms of iteration complexity.

Citations

Cited by

Related