Constant Step Size Stochastic Gradient Descent for Probabilistic Modeling

Dmitry Babichev, Francis Bach
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:218-227, 2018.

Abstract

Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with param- eter averaging, these are not leading to con- vergent algorithms in general. In this pa- per, we consider generalized linear models, that is, conditional models based on exponen- tial families. We propose averaging moment parameters instead of natural parameters for constant-step-size stochastic gradient descent. For finite-dimensional models, we show that this can sometimes (and surprisingly) lead to better predictions than the best linear model. For infinite-dimensional models, we show that it always converges to optimal predictions, while averaging natural parameters never does. We illustrate our findings with simulations on synthetic data and classical benchmarks with many observations.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-babichev18a, title = {Constant Step Size Stochastic Gradient Descent for Probabilistic Modeling}, author = {Babichev, Dmitry and Bach, Francis}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {218--227}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/babichev18a/babichev18a.pdf}, url = {https://proceedings.mlr.press/r16/babichev18a.html}, abstract = {Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with param- eter averaging, these are not leading to con- vergent algorithms in general. In this pa- per, we consider generalized linear models, that is, conditional models based on exponen- tial families. We propose averaging moment parameters instead of natural parameters for constant-step-size stochastic gradient descent. For finite-dimensional models, we show that this can sometimes (and surprisingly) lead to better predictions than the best linear model. For infinite-dimensional models, we show that it always converges to optimal predictions, while averaging natural parameters never does. We illustrate our findings with simulations on synthetic data and classical benchmarks with many observations.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Constant Step Size Stochastic Gradient Descent for Probabilistic Modeling %A Dmitry Babichev %A Francis Bach %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-babichev18a %I PMLR %P 218--227 %U https://proceedings.mlr.press/r16/babichev18a.html %V R16 %X Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with param- eter averaging, these are not leading to con- vergent algorithms in general. In this pa- per, we consider generalized linear models, that is, conditional models based on exponen- tial families. We propose averaging moment parameters instead of natural parameters for constant-step-size stochastic gradient descent. For finite-dimensional models, we show that this can sometimes (and surprisingly) lead to better predictions than the best linear model. For infinite-dimensional models, we show that it always converges to optimal predictions, while averaging natural parameters never does. We illustrate our findings with simulations on synthetic data and classical benchmarks with many observations. %Z Reissued by PMLR on 04 October 2026.
APA
Babichev, D. & Bach, F.. (2018). Constant Step Size Stochastic Gradient Descent for Probabilistic Modeling. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:218-227 Available from https://proceedings.mlr.press/r16/babichev18a.html. Reissued by PMLR on 04 October 2026.

Related Material