[edit]
Coupling Adaptive Batch Sizes with Learning Rates
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:241-250, 2017.
Abstract
Mini-batch stochastic gradient descent and variants thereof have become standard for large-scale empirical risk minimization like the training of neural networks. These meth- ods are usually used with a constant batch size chosen by simple empirical inspection. The batch size significantly influences the behav- ior of the stochastic optimization algorithm, though, since it determines the variance of the gradient estimates. This variance also changes over the optimization process; when using a constant batch size, stability and convergence is thus often enforced by means of a (manually tuned) decreasing learning rate schedule. We propose a practical method for dynamic batch size adaptation. It estimates the vari- ance of the stochastic gradients and adapts the batch size to decrease the variance proportion- ally to the value of the objective function, re- moving the need for the aforementioned learn- ing rate decrease. In contrast to recent related work, our algorithm couples the batch size to the learning rate, directly reflecting the known relationship between the two. On popular im- age classification benchmarks, our batch size adaptation yields faster optimization conver- gence, while simultaneously simplifying learn- ing rate tuning. A TensorFlow implementation is available.