Reassessing Muon for Matrix Factorization

Ali Parviz, Gal Mishne, Alex Cloninger
Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026), PMLR 334(2):59-86, 2026.

Abstract

Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approx- imate orthogonalization and has been reported to outperform Adam and AdamW on large language model training. Its empirical success has motivated a growing theoretical liter- ature that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon{’}s advantages stem from its update rule itself and which are arti- facts of the scale, architecture, and data of modern deep networks. In this work we isolate the optimizer from these confounders by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled and systematically tuned comparison against adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting, and that several previously reported advan- tages are sensitive to hyperparameter choices. Our results give a more nuanced picture of when spectrum-aware orthogonalization helps, and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v334-parviz26a, title = {Reassessing Muon for Matrix Factorization}, author = {Parviz, Ali and Mishne, Gal and Cloninger, Alex}, booktitle = {Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026)}, pages = {59--86}, year = {2026}, editor = {Berman, Eddie and Bernárdez, Guillermo and Chen, Samantha and Cloninger, Alex and Doster, Timothy and Emerson, Tegan and Grigsby, J. Elisenda and Kvinge, Henry and Lawrence, Hannah and Marrinan, Tim and Myers, Audun and Papillon, Mathilde and Tahmasebi, Behrooz and Telyatnikov, Lev and Walters, Robin and Weber, Melanie and Xie, YuQing and Yeats, Eric}, volume = {334}, number = {2}, series = {Proceedings of Machine Learning Research}, month = {18--20 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v334/main/assets/parviz26a/parviz26a.pdf}, url = {https://proceedings.mlr.press/v334/parviz26a.html}, abstract = {Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approx- imate orthogonalization and has been reported to outperform Adam and AdamW on large language model training. Its empirical success has motivated a growing theoretical liter- ature that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon{’}s advantages stem from its update rule itself and which are arti- facts of the scale, architecture, and data of modern deep networks. In this work we isolate the optimizer from these confounders by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled and systematically tuned comparison against adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting, and that several previously reported advan- tages are sensitive to hyperparameter choices. Our results give a more nuanced picture of when spectrum-aware orthogonalization helps, and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.} }
Endnote
%0 Conference Paper %T Reassessing Muon for Matrix Factorization %A Ali Parviz %A Gal Mishne %A Alex Cloninger %B Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026) %C Proceedings of Machine Learning Research %D 2026 %E Eddie Berman %E Guillermo Bernárdez %E Samantha Chen %E Alex Cloninger %E Timothy Doster %E Tegan Emerson %E J. Elisenda Grigsby %E Henry Kvinge %E Hannah Lawrence %E Tim Marrinan %E Audun Myers %E Mathilde Papillon %E Behrooz Tahmasebi %E Lev Telyatnikov %E Robin Walters %E Melanie Weber %E YuQing Xie %E Eric Yeats %F pmlr-v334-parviz26a %I PMLR %P 59--86 %U https://proceedings.mlr.press/v334/parviz26a.html %V 334 %N 2 %X Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approx- imate orthogonalization and has been reported to outperform Adam and AdamW on large language model training. Its empirical success has motivated a growing theoretical liter- ature that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon{’}s advantages stem from its update rule itself and which are arti- facts of the scale, architecture, and data of modern deep networks. In this work we isolate the optimizer from these confounders by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled and systematically tuned comparison against adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting, and that several previously reported advan- tages are sensitive to hyperparameter choices. Our results give a more nuanced picture of when spectrum-aware orthogonalization helps, and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.
APA
Parviz, A., Mishne, G. & Cloninger, A.. (2026). Reassessing Muon for Matrix Factorization. Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026), in Proceedings of Machine Learning Research 334(2):59-86 Available from https://proceedings.mlr.press/v334/parviz26a.html.

Related Material