[edit]
Stable Gradient Descent
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:765-774, 2018.
Abstract
The goal of many machine learning tasks is to learn a model that has small population risk. While mini-batch stochastic gradient descent (SGD) and variants are popular approaches for achieving this goal, it is hard to pre- scribe a clear stopping criterion and to establish high probability convergence bounds to the population risk. In this paper, we introduce Stable Gradient Descent which validates stochastic gra- dient computations by splitting data into training and validation sets and reuses samples using a differential pri- vate mechanism. StGD comes with a natural upper bound on the number of iterations and has high-probability convergence to the population risk. Ex- perimental results illustrate that StGD is empirically competitive and often better than SGD and GD.