Averaging Weights Leads to Wider Optima and Better Generalization

Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, Andrew Gordon Wilson
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:875-884, 2018.

Abstract

Deep neural networks are typically trained by optimizing a loss function with an SGD vari- ant, in conjunction with a decaying learning rate, until convergence. We show that simple averaging of multiple points along the trajec- tory of SGD, with a cyclical or constant learn- ing rate, leads to better generalization than conventional training. We also show that this Stochastic Weight Averaging (SWA) procedure finds much broader optima than SGD, and ap- proximates the recent Fast Geometric Ensem- bling (FGE) approach with a single model. Using SWA we achieve notable improvement in test accuracy over conventional SGD train- ing on a range of state-of-the-art residual net- works, PyramidNets, DenseNets, and Shake- Shake networks on CIFAR-10, CIFAR-100, and ImageNet. In short, SWA is extremely easy to implement, improves generalization, and has almost no computational overhead.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-izmailov18a, title = {Averaging Weights Leads to Wider Optima and Better Generalization}, author = {Izmailov, Pavel and Podoprikhin, Dmitrii and Garipov, Timur and Vetrov, Dmitry and Wilson, Andrew Gordon}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {875--884}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/izmailov18a/izmailov18a.pdf}, url = {https://proceedings.mlr.press/r16/izmailov18a.html}, abstract = {Deep neural networks are typically trained by optimizing a loss function with an SGD vari- ant, in conjunction with a decaying learning rate, until convergence. We show that simple averaging of multiple points along the trajec- tory of SGD, with a cyclical or constant learn- ing rate, leads to better generalization than conventional training. We also show that this Stochastic Weight Averaging (SWA) procedure finds much broader optima than SGD, and ap- proximates the recent Fast Geometric Ensem- bling (FGE) approach with a single model. Using SWA we achieve notable improvement in test accuracy over conventional SGD train- ing on a range of state-of-the-art residual net- works, PyramidNets, DenseNets, and Shake- Shake networks on CIFAR-10, CIFAR-100, and ImageNet. In short, SWA is extremely easy to implement, improves generalization, and has almost no computational overhead.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Averaging Weights Leads to Wider Optima and Better Generalization %A Pavel Izmailov %A Dmitrii Podoprikhin %A Timur Garipov %A Dmitry Vetrov %A Andrew Gordon Wilson %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-izmailov18a %I PMLR %P 875--884 %U https://proceedings.mlr.press/r16/izmailov18a.html %V R16 %X Deep neural networks are typically trained by optimizing a loss function with an SGD vari- ant, in conjunction with a decaying learning rate, until convergence. We show that simple averaging of multiple points along the trajec- tory of SGD, with a cyclical or constant learn- ing rate, leads to better generalization than conventional training. We also show that this Stochastic Weight Averaging (SWA) procedure finds much broader optima than SGD, and ap- proximates the recent Fast Geometric Ensem- bling (FGE) approach with a single model. Using SWA we achieve notable improvement in test accuracy over conventional SGD train- ing on a range of state-of-the-art residual net- works, PyramidNets, DenseNets, and Shake- Shake networks on CIFAR-10, CIFAR-100, and ImageNet. In short, SWA is extremely easy to implement, improves generalization, and has almost no computational overhead. %Z Reissued by PMLR on 04 October 2026.
APA
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D. & Wilson, A.G.. (2018). Averaging Weights Leads to Wider Optima and Better Generalization. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:875-884 Available from https://proceedings.mlr.press/r16/izmailov18a.html. Reissued by PMLR on 04 October 2026.

Related Material