Coupling Adaptive Batch Sizes with Learning Rates

Lukas Balles, Javier Romero, Philipp Hennig
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:241-250, 2017.

Abstract

Mini-batch stochastic gradient descent and variants thereof have become standard for large-scale empirical risk minimization like the training of neural networks. These meth- ods are usually used with a constant batch size chosen by simple empirical inspection. The batch size significantly influences the behav- ior of the stochastic optimization algorithm, though, since it determines the variance of the gradient estimates. This variance also changes over the optimization process; when using a constant batch size, stability and convergence is thus often enforced by means of a (manually tuned) decreasing learning rate schedule. We propose a practical method for dynamic batch size adaptation. It estimates the vari- ance of the stochastic gradients and adapts the batch size to decrease the variance proportion- ally to the value of the objective function, re- moving the need for the aforementioned learn- ing rate decrease. In contrast to recent related work, our algorithm couples the batch size to the learning rate, directly reflecting the known relationship between the two. On popular im- age classification benchmarks, our batch size adaptation yields faster optimization conver- gence, while simultaneously simplifying learn- ing rate tuning. A TensorFlow implementation is available.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR15-balles17a, title = {Coupling Adaptive Batch Sizes with Learning Rates}, author = {Balles, Lukas and Romero, Javier and Hennig, Philipp}, booktitle = {Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence}, pages = {241--250}, year = {2017}, editor = {Elidan, Gal and Kersting, Kristian}, volume = {R15}, series = {Proceedings of Machine Learning Research}, month = {11--15 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r15/main/assets/balles17a/balles17a.pdf}, url = {https://proceedings.mlr.press/r15/balles17a.html}, abstract = {Mini-batch stochastic gradient descent and variants thereof have become standard for large-scale empirical risk minimization like the training of neural networks. These meth- ods are usually used with a constant batch size chosen by simple empirical inspection. The batch size significantly influences the behav- ior of the stochastic optimization algorithm, though, since it determines the variance of the gradient estimates. This variance also changes over the optimization process; when using a constant batch size, stability and convergence is thus often enforced by means of a (manually tuned) decreasing learning rate schedule. We propose a practical method for dynamic batch size adaptation. It estimates the vari- ance of the stochastic gradients and adapts the batch size to decrease the variance proportion- ally to the value of the objective function, re- moving the need for the aforementioned learn- ing rate decrease. In contrast to recent related work, our algorithm couples the batch size to the learning rate, directly reflecting the known relationship between the two. On popular im- age classification benchmarks, our batch size adaptation yields faster optimization conver- gence, while simultaneously simplifying learn- ing rate tuning. A TensorFlow implementation is available.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Coupling Adaptive Batch Sizes with Learning Rates %A Lukas Balles %A Javier Romero %A Philipp Hennig %B Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2017 %E Gal Elidan %E Kristian Kersting %F pmlr-vR15-balles17a %I PMLR %P 241--250 %U https://proceedings.mlr.press/r15/balles17a.html %V R15 %X Mini-batch stochastic gradient descent and variants thereof have become standard for large-scale empirical risk minimization like the training of neural networks. These meth- ods are usually used with a constant batch size chosen by simple empirical inspection. The batch size significantly influences the behav- ior of the stochastic optimization algorithm, though, since it determines the variance of the gradient estimates. This variance also changes over the optimization process; when using a constant batch size, stability and convergence is thus often enforced by means of a (manually tuned) decreasing learning rate schedule. We propose a practical method for dynamic batch size adaptation. It estimates the vari- ance of the stochastic gradients and adapts the batch size to decrease the variance proportion- ally to the value of the objective function, re- moving the need for the aforementioned learn- ing rate decrease. In contrast to recent related work, our algorithm couples the batch size to the learning rate, directly reflecting the known relationship between the two. On popular im- age classification benchmarks, our batch size adaptation yields faster optimization conver- gence, while simultaneously simplifying learn- ing rate tuning. A TensorFlow implementation is available. %Z Reissued by PMLR on 04 October 2026.
APA
Balles, L., Romero, J. & Hennig, P.. (2017). Coupling Adaptive Batch Sizes with Learning Rates. Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R15:241-250 Available from https://proceedings.mlr.press/r15/balles17a.html. Reissued by PMLR on 04 October 2026.

Related Material