[edit]
Towards Flatter Loss Surface via Nonmonotonic Learning Rate Scheduling
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:1019-1029, 2018.
Abstract
Whereas optimizing deep neural networks us- ing stochastic gradient descent has shown great performances in practice, the rule for setting step size (i.e. learning rate) of gradient de- scent is not well studied. Although it appears that some intriguing learning rate rules such as ADAM (Kingma and Ba, 2014) have since been developed, they concentrated on improv- ing convergence, not on improving generaliza- tion capabilities. Recently, the improved gen- eralization property of the flat minima was re- visited, and this research guides us towards promising solutions to many current optimiza- tion problems. In this paper, we analyze the flatness of loss surfaces through the lens of ro- bustness to input perturbations and advocate that gradient descent should be guided to reach flatter region of loss surfaces to achieve gen- eralization. Finally, we suggest a learning rate rule for escaping sharp regions of loss surfaces, and we demonstrate the capacity of our ap- proach by performing numerous experiments.