Stable Gradient Descent

Yingxue Zhou, Sheng Chen, Arindam Banerjee
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:765-774, 2018.

Abstract

The goal of many machine learning tasks is to learn a model that has small population risk. While mini-batch stochastic gradient descent (SGD) and variants are popular approaches for achieving this goal, it is hard to pre- scribe a clear stopping criterion and to establish high probability convergence bounds to the population risk. In this paper, we introduce Stable Gradient Descent which validates stochastic gra- dient computations by splitting data into training and validation sets and reuses samples using a differential pri- vate mechanism. StGD comes with a natural upper bound on the number of iterations and has high-probability convergence to the population risk. Ex- perimental results illustrate that StGD is empirically competitive and often better than SGD and GD.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-zhou18a, title = {Stable Gradient Descent}, author = {Zhou, Yingxue and Chen, Sheng and Banerjee, Arindam}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {765--774}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/zhou18a/zhou18a.pdf}, url = {https://proceedings.mlr.press/r16/zhou18a.html}, abstract = {The goal of many machine learning tasks is to learn a model that has small population risk. While mini-batch stochastic gradient descent (SGD) and variants are popular approaches for achieving this goal, it is hard to pre- scribe a clear stopping criterion and to establish high probability convergence bounds to the population risk. In this paper, we introduce Stable Gradient Descent which validates stochastic gra- dient computations by splitting data into training and validation sets and reuses samples using a differential pri- vate mechanism. StGD comes with a natural upper bound on the number of iterations and has high-probability convergence to the population risk. Ex- perimental results illustrate that StGD is empirically competitive and often better than SGD and GD.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Stable Gradient Descent %A Yingxue Zhou %A Sheng Chen %A Arindam Banerjee %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-zhou18a %I PMLR %P 765--774 %U https://proceedings.mlr.press/r16/zhou18a.html %V R16 %X The goal of many machine learning tasks is to learn a model that has small population risk. While mini-batch stochastic gradient descent (SGD) and variants are popular approaches for achieving this goal, it is hard to pre- scribe a clear stopping criterion and to establish high probability convergence bounds to the population risk. In this paper, we introduce Stable Gradient Descent which validates stochastic gra- dient computations by splitting data into training and validation sets and reuses samples using a differential pri- vate mechanism. StGD comes with a natural upper bound on the number of iterations and has high-probability convergence to the population risk. Ex- perimental results illustrate that StGD is empirically competitive and often better than SGD and GD. %Z Reissued by PMLR on 04 October 2026.
APA
Zhou, Y., Chen, S. & Banerjee, A.. (2018). Stable Gradient Descent. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:765-774 Available from https://proceedings.mlr.press/r16/zhou18a.html. Reissued by PMLR on 04 October 2026.

Related Material