Towards Flatter Loss Surface via Nonmonotonic Learning Rate Scheduling

Sihyeon Seong, Yegang Lee, Youngwook Kee, Dongyoon Han, Junmo Kim
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:1019-1029, 2018.

Abstract

Whereas optimizing deep neural networks us- ing stochastic gradient descent has shown great performances in practice, the rule for setting step size (i.e. learning rate) of gradient de- scent is not well studied. Although it appears that some intriguing learning rate rules such as ADAM (Kingma and Ba, 2014) have since been developed, they concentrated on improv- ing convergence, not on improving generaliza- tion capabilities. Recently, the improved gen- eralization property of the flat minima was re- visited, and this research guides us towards promising solutions to many current optimiza- tion problems. In this paper, we analyze the flatness of loss surfaces through the lens of ro- bustness to input perturbations and advocate that gradient descent should be guided to reach flatter region of loss surfaces to achieve gen- eralization. Finally, we suggest a learning rate rule for escaping sharp regions of loss surfaces, and we demonstrate the capacity of our ap- proach by performing numerous experiments.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-seong18a, title = {Towards Flatter Loss Surface via Nonmonotonic Learning Rate Scheduling}, author = {Seong, Sihyeon and Lee, Yegang and Kee, Youngwook and Han, Dongyoon and Kim, Junmo}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {1019--1029}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/seong18a/seong18a.pdf}, url = {https://proceedings.mlr.press/r16/seong18a.html}, abstract = {Whereas optimizing deep neural networks us- ing stochastic gradient descent has shown great performances in practice, the rule for setting step size (i.e. learning rate) of gradient de- scent is not well studied. Although it appears that some intriguing learning rate rules such as ADAM (Kingma and Ba, 2014) have since been developed, they concentrated on improv- ing convergence, not on improving generaliza- tion capabilities. Recently, the improved gen- eralization property of the flat minima was re- visited, and this research guides us towards promising solutions to many current optimiza- tion problems. In this paper, we analyze the flatness of loss surfaces through the lens of ro- bustness to input perturbations and advocate that gradient descent should be guided to reach flatter region of loss surfaces to achieve gen- eralization. Finally, we suggest a learning rate rule for escaping sharp regions of loss surfaces, and we demonstrate the capacity of our ap- proach by performing numerous experiments.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Towards Flatter Loss Surface via Nonmonotonic Learning Rate Scheduling %A Sihyeon Seong %A Yegang Lee %A Youngwook Kee %A Dongyoon Han %A Junmo Kim %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-seong18a %I PMLR %P 1019--1029 %U https://proceedings.mlr.press/r16/seong18a.html %V R16 %X Whereas optimizing deep neural networks us- ing stochastic gradient descent has shown great performances in practice, the rule for setting step size (i.e. learning rate) of gradient de- scent is not well studied. Although it appears that some intriguing learning rate rules such as ADAM (Kingma and Ba, 2014) have since been developed, they concentrated on improv- ing convergence, not on improving generaliza- tion capabilities. Recently, the improved gen- eralization property of the flat minima was re- visited, and this research guides us towards promising solutions to many current optimiza- tion problems. In this paper, we analyze the flatness of loss surfaces through the lens of ro- bustness to input perturbations and advocate that gradient descent should be guided to reach flatter region of loss surfaces to achieve gen- eralization. Finally, we suggest a learning rate rule for escaping sharp regions of loss surfaces, and we demonstrate the capacity of our ap- proach by performing numerous experiments. %Z Reissued by PMLR on 04 October 2026.
APA
Seong, S., Lee, Y., Kee, Y., Han, D. & Kim, J.. (2018). Towards Flatter Loss Surface via Nonmonotonic Learning Rate Scheduling. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:1019-1029 Available from https://proceedings.mlr.press/r16/seong18a.html. Reissued by PMLR on 04 October 2026.

Related Material