Understanding SAM through Minimax Perspective

Ying Chen, Aoxi Li, Javad Lavaei
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:15170-15193, 2026.

Abstract

Sharpness-Aware Minimization (SAM) empirically boosts generalization by seeking parameters that minimize the worst-case loss in a small neighborhood, yet existing theory explains its behavior under either Polyak-Lojasiewicz (PL) condition or upper bounded perturbation radius. We revisit SAM through the bilevel minimax problem $\min_{\theta}\max_{\|\Delta\|\le\rho}l(\theta+\Delta)$ and derive a $(\theta,\Delta)$ gradient flow ODE whose equilibria coincide with the problem’s optimality conditions. A Lyapunov argument-free of convexity assumptions, quantifies how the optimality gap depends on the radius $\rho$ and local curvature. Discretizing the flow yields a Multi-step SAM algorithm that recovers classical SAM as $\rho\to 0$. Moreover, our analysis and the resulting algorithm remain valid even for large $\rho$, providing guidance for aggressive neighborhood exploration. Experiments on synthetic objectives and CIFAR-10 validate the predicted gains from multiple inner updates, bridging the gap between SAM’s minimax intuition and its practical implementation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26bo, title = {Understanding {SAM} through Minimax Perspective}, author = {Chen, Ying and Li, Aoxi and Lavaei, Javad}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {15170--15193}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26bo/chen26bo.pdf}, url = {https://proceedings.mlr.press/v306/chen26bo.html}, abstract = {Sharpness-Aware Minimization (SAM) empirically boosts generalization by seeking parameters that minimize the worst-case loss in a small neighborhood, yet existing theory explains its behavior under either Polyak-Lojasiewicz (PL) condition or upper bounded perturbation radius. We revisit SAM through the bilevel minimax problem $\min_{\theta}\max_{\|\Delta\|\le\rho}l(\theta+\Delta)$ and derive a $(\theta,\Delta)$ gradient flow ODE whose equilibria coincide with the problem’s optimality conditions. A Lyapunov argument-free of convexity assumptions, quantifies how the optimality gap depends on the radius $\rho$ and local curvature. Discretizing the flow yields a Multi-step SAM algorithm that recovers classical SAM as $\rho\to 0$. Moreover, our analysis and the resulting algorithm remain valid even for large $\rho$, providing guidance for aggressive neighborhood exploration. Experiments on synthetic objectives and CIFAR-10 validate the predicted gains from multiple inner updates, bridging the gap between SAM’s minimax intuition and its practical implementation.} }
Endnote
%0 Conference Paper %T Understanding SAM through Minimax Perspective %A Ying Chen %A Aoxi Li %A Javad Lavaei %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26bo %I PMLR %P 15170--15193 %U https://proceedings.mlr.press/v306/chen26bo.html %V 306 %X Sharpness-Aware Minimization (SAM) empirically boosts generalization by seeking parameters that minimize the worst-case loss in a small neighborhood, yet existing theory explains its behavior under either Polyak-Lojasiewicz (PL) condition or upper bounded perturbation radius. We revisit SAM through the bilevel minimax problem $\min_{\theta}\max_{\|\Delta\|\le\rho}l(\theta+\Delta)$ and derive a $(\theta,\Delta)$ gradient flow ODE whose equilibria coincide with the problem’s optimality conditions. A Lyapunov argument-free of convexity assumptions, quantifies how the optimality gap depends on the radius $\rho$ and local curvature. Discretizing the flow yields a Multi-step SAM algorithm that recovers classical SAM as $\rho\to 0$. Moreover, our analysis and the resulting algorithm remain valid even for large $\rho$, providing guidance for aggressive neighborhood exploration. Experiments on synthetic objectives and CIFAR-10 validate the predicted gains from multiple inner updates, bridging the gap between SAM’s minimax intuition and its practical implementation.
APA
Chen, Y., Li, A. & Lavaei, J.. (2026). Understanding SAM through Minimax Perspective. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:15170-15193 Available from https://proceedings.mlr.press/v306/chen26bo.html.

Related Material