Exploiting weight-space symmetries for approximating curvature

Artem Artemev, Rui Xia, Benjamin M. Boyd, Youjing Yu, Felix Dangel, Guillaume Hennequin, Alberto Bernacchia
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:3856-3890, 2026.

Abstract

Many machine learning techniques rely on approximating a loss function’s curvature, but this is notoriously hard to do at the scale of modern deep networks. Surprisingly, no previous work has exploited the curvature constraints that arise from well known weight-space symmetries in loss landscapes. By analytically averaging over group actions that leave the loss invariant, we construct structured Hessian approximations from single gradients that can be tractably estimated, stored, and inverted. The choice of user-specified symmetry group directly governs the trade-off between approximation accuracy and computational cost. Moreover, our framework provides a unifying theoretical lens for viewing existing methods; in particular, a specific choice of symmetry group recovers Shampoo/Muon-like curvature estimates. We validate our method on a range of network architectures, and deploy it to second-order optimization benchmarks, including a small language model. Our curvature estimation framework might find applications in other machine learning problems such as uncertainty estimation, continual learning, compression/pruning, training data attribution, and more.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-artemev26a, title = {Exploiting weight-space symmetries for approximating curvature}, author = {Artemev, Artem and Xia, Rui and Boyd, Benjamin M. and Yu, Youjing and Dangel, Felix and Hennequin, Guillaume and Bernacchia, Alberto}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {3856--3890}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/artemev26a/artemev26a.pdf}, url = {https://proceedings.mlr.press/v306/artemev26a.html}, abstract = {Many machine learning techniques rely on approximating a loss function’s curvature, but this is notoriously hard to do at the scale of modern deep networks. Surprisingly, no previous work has exploited the curvature constraints that arise from well known weight-space symmetries in loss landscapes. By analytically averaging over group actions that leave the loss invariant, we construct structured Hessian approximations from single gradients that can be tractably estimated, stored, and inverted. The choice of user-specified symmetry group directly governs the trade-off between approximation accuracy and computational cost. Moreover, our framework provides a unifying theoretical lens for viewing existing methods; in particular, a specific choice of symmetry group recovers Shampoo/Muon-like curvature estimates. We validate our method on a range of network architectures, and deploy it to second-order optimization benchmarks, including a small language model. Our curvature estimation framework might find applications in other machine learning problems such as uncertainty estimation, continual learning, compression/pruning, training data attribution, and more.} }
Endnote
%0 Conference Paper %T Exploiting weight-space symmetries for approximating curvature %A Artem Artemev %A Rui Xia %A Benjamin M. Boyd %A Youjing Yu %A Felix Dangel %A Guillaume Hennequin %A Alberto Bernacchia %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-artemev26a %I PMLR %P 3856--3890 %U https://proceedings.mlr.press/v306/artemev26a.html %V 306 %X Many machine learning techniques rely on approximating a loss function’s curvature, but this is notoriously hard to do at the scale of modern deep networks. Surprisingly, no previous work has exploited the curvature constraints that arise from well known weight-space symmetries in loss landscapes. By analytically averaging over group actions that leave the loss invariant, we construct structured Hessian approximations from single gradients that can be tractably estimated, stored, and inverted. The choice of user-specified symmetry group directly governs the trade-off between approximation accuracy and computational cost. Moreover, our framework provides a unifying theoretical lens for viewing existing methods; in particular, a specific choice of symmetry group recovers Shampoo/Muon-like curvature estimates. We validate our method on a range of network architectures, and deploy it to second-order optimization benchmarks, including a small language model. Our curvature estimation framework might find applications in other machine learning problems such as uncertainty estimation, continual learning, compression/pruning, training data attribution, and more.
APA
Artemev, A., Xia, R., Boyd, B.M., Yu, Y., Dangel, F., Hennequin, G. & Bernacchia, A.. (2026). Exploiting weight-space symmetries for approximating curvature. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:3856-3890 Available from https://proceedings.mlr.press/v306/artemev26a.html.

Related Material