Annealed Gradient Descent for Deep Learning

Hengyue Pan York University, Hui Jiang York University
Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, PMLR R13:189-198, 2015.

Abstract

In this paper, we propose a novel annealed gradient descent (AGD) method for deep learning. AGD optimizes a sequence of gradually improved smoother mosaic functions that approximate the original non-convex objective function according to an annealing schedule during optimization process. We present a theoretical analysis on its convergence properties and learning speed. The proposed AGD algorithm is applied to learning deep neural networks (DNN) for image recognition in MNIST and speech recognition in Switchboard. Experimental results have shown that AGD can yield comparable performance as SGD but it can significantly expedite training of DNNs in big data sets (by about 40% faster).

Cite this Paper


BibTeX
@InProceedings{pmlr-vR13-university15c, title = {Annealed Gradient Descent for Deep Learning}, author = {University, Hengyue Pan York and University, Hui Jiang York}, booktitle = {Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence}, pages = {189--198}, year = {2015}, editor = {Meila, Marina and Heskes, Tom}, volume = {R13}, series = {Proceedings of Machine Learning Research}, month = {12--16 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r13/main/assets/university15c/university15c.pdf}, url = {https://proceedings.mlr.press/r13/university15c.html}, abstract = {In this paper, we propose a novel annealed gradient descent (AGD) method for deep learning. AGD optimizes a sequence of gradually improved smoother mosaic functions that approximate the original non-convex objective function according to an annealing schedule during optimization process. We present a theoretical analysis on its convergence properties and learning speed. The proposed AGD algorithm is applied to learning deep neural networks (DNN) for image recognition in MNIST and speech recognition in Switchboard. Experimental results have shown that AGD can yield comparable performance as SGD but it can significantly expedite training of DNNs in big data sets (by about 40% faster).}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Annealed Gradient Descent for Deep Learning %A Hengyue Pan York University %A Hui Jiang York University %B Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2015 %E Marina Meila %E Tom Heskes %F pmlr-vR13-university15c %I PMLR %P 189--198 %U https://proceedings.mlr.press/r13/university15c.html %V R13 %X In this paper, we propose a novel annealed gradient descent (AGD) method for deep learning. AGD optimizes a sequence of gradually improved smoother mosaic functions that approximate the original non-convex objective function according to an annealing schedule during optimization process. We present a theoretical analysis on its convergence properties and learning speed. The proposed AGD algorithm is applied to learning deep neural networks (DNN) for image recognition in MNIST and speech recognition in Switchboard. Experimental results have shown that AGD can yield comparable performance as SGD but it can significantly expedite training of DNNs in big data sets (by about 40% faster). %Z Reissued by PMLR on 04 October 2026.
APA
University, H.P.Y. & University, H.J.Y.. (2015). Annealed Gradient Descent for Deep Learning. Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R13:189-198 Available from https://proceedings.mlr.press/r13/university15c.html. Reissued by PMLR on 04 October 2026.

Related Material