An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants

Michael Crawshaw, Chirag Modi, Mingrui Liu, Robert M. Gower
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:21681-21716, 2026.

Abstract

To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization. We systematically explore different alternatives for aggregating norms across layers, both formalizing existing combinations of Adam and the recently proposed Muon as a type of non-Euclidean gradient descent, and deriving new variants of the Muon optimizer. Through a comprehensive experimental evaluation of the optimizers within our framework, we find that Muon is sensitive to the choice of learning rate, whereas a new variant we call MuonMax is significantly more robust. We then show how to combine any non-Euclidean gradient method with model based momentum (known as Momo). The new Momo variants of Muon are significantly more robust to hyperparameter tuning, and often achieve a better validation score. Thus for new tasks, where the optimal hyperparameters are not known, we advocate for using Momo in combination with MuonMax to save on costly hyperparameter tuning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-crawshaw26a, title = {An Exploration of Non-{E}uclidean Gradient Descent: Muon and its Many Variants}, author = {Crawshaw, Michael and Modi, Chirag and Liu, Mingrui and Gower, Robert M.}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {21681--21716}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/crawshaw26a/crawshaw26a.pdf}, url = {https://proceedings.mlr.press/v306/crawshaw26a.html}, abstract = {To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization. We systematically explore different alternatives for aggregating norms across layers, both formalizing existing combinations of Adam and the recently proposed Muon as a type of non-Euclidean gradient descent, and deriving new variants of the Muon optimizer. Through a comprehensive experimental evaluation of the optimizers within our framework, we find that Muon is sensitive to the choice of learning rate, whereas a new variant we call MuonMax is significantly more robust. We then show how to combine any non-Euclidean gradient method with model based momentum (known as Momo). The new Momo variants of Muon are significantly more robust to hyperparameter tuning, and often achieve a better validation score. Thus for new tasks, where the optimal hyperparameters are not known, we advocate for using Momo in combination with MuonMax to save on costly hyperparameter tuning.} }
Endnote
%0 Conference Paper %T An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants %A Michael Crawshaw %A Chirag Modi %A Mingrui Liu %A Robert M. Gower %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-crawshaw26a %I PMLR %P 21681--21716 %U https://proceedings.mlr.press/v306/crawshaw26a.html %V 306 %X To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization. We systematically explore different alternatives for aggregating norms across layers, both formalizing existing combinations of Adam and the recently proposed Muon as a type of non-Euclidean gradient descent, and deriving new variants of the Muon optimizer. Through a comprehensive experimental evaluation of the optimizers within our framework, we find that Muon is sensitive to the choice of learning rate, whereas a new variant we call MuonMax is significantly more robust. We then show how to combine any non-Euclidean gradient method with model based momentum (known as Momo). The new Momo variants of Muon are significantly more robust to hyperparameter tuning, and often achieve a better validation score. Thus for new tasks, where the optimal hyperparameters are not known, we advocate for using Momo in combination with MuonMax to save on costly hyperparameter tuning.
APA
Crawshaw, M., Modi, C., Liu, M. & Gower, R.M.. (2026). An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:21681-21716 Available from https://proceedings.mlr.press/v306/crawshaw26a.html.

Related Material