[edit]
Averaging Weights Leads to Wider Optima and Better Generalization
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:875-884, 2018.
Abstract
Deep neural networks are typically trained by optimizing a loss function with an SGD vari- ant, in conjunction with a decaying learning rate, until convergence. We show that simple averaging of multiple points along the trajec- tory of SGD, with a cyclical or constant learn- ing rate, leads to better generalization than conventional training. We also show that this Stochastic Weight Averaging (SWA) procedure finds much broader optima than SGD, and ap- proximates the recent Fast Geometric Ensem- bling (FGE) approach with a single model. Using SWA we achieve notable improvement in test accuracy over conventional SGD train- ing on a range of state-of-the-art residual net- works, PyramidNets, DenseNets, and Shake- Shake networks on CIFAR-10, CIFAR-100, and ImageNet. In short, SWA is extremely easy to implement, improves generalization, and has almost no computational overhead.