General Weighted Averaging in Stochastic Gradient Descent: CLT and Adaptive Optimality

Ziyang Wei, Wanrong Zhu, Wei Biao Wu
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1198-1206, 2026.

Abstract

Stochastic Gradient Descent (SGD) is a cornerstone of machine learning, prized for its efficiency in large-scale optimization. This paper revisits SGD by introducing a general weighted averaging framework that significantly enhances its applicability. We establish asymptotic normality for a wide range of weighted averaged SGD solutions under minimal assumptions, providing a groundbreaking necessary condition for the central limit theorem in certain settings. This enables asymptotically valid online inference, empowering real-time confidence interval construction. Furthermore, we propose an adaptive averaging scheme, inspired by optimal weights for linear models, which achieves optimal superior non-asymptotic bounds. Our theoretical advances and empirical validations redefine SGD’s capabilities, offering transformative insights for statistical learning and optimization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-wei26b, title = { General Weighted Averaging in Stochastic Gradient Descent: CLT and Adaptive Optimality }, author = {Wei, Ziyang and Zhu, Wanrong and Wu, Wei Biao}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1198--1206}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/wei26b/wei26b.pdf}, url = {https://proceedings.mlr.press/v300/wei26b.html}, abstract = { Stochastic Gradient Descent (SGD) is a cornerstone of machine learning, prized for its efficiency in large-scale optimization. This paper revisits SGD by introducing a general weighted averaging framework that significantly enhances its applicability. We establish asymptotic normality for a wide range of weighted averaged SGD solutions under minimal assumptions, providing a groundbreaking necessary condition for the central limit theorem in certain settings. This enables asymptotically valid online inference, empowering real-time confidence interval construction. Furthermore, we propose an adaptive averaging scheme, inspired by optimal weights for linear models, which achieves optimal superior non-asymptotic bounds. Our theoretical advances and empirical validations redefine SGD’s capabilities, offering transformative insights for statistical learning and optimization. } }
Endnote
%0 Conference Paper %T General Weighted Averaging in Stochastic Gradient Descent: CLT and Adaptive Optimality %A Ziyang Wei %A Wanrong Zhu %A Wei Biao Wu %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-wei26b %I PMLR %P 1198--1206 %U https://proceedings.mlr.press/v300/wei26b.html %V 300 %X Stochastic Gradient Descent (SGD) is a cornerstone of machine learning, prized for its efficiency in large-scale optimization. This paper revisits SGD by introducing a general weighted averaging framework that significantly enhances its applicability. We establish asymptotic normality for a wide range of weighted averaged SGD solutions under minimal assumptions, providing a groundbreaking necessary condition for the central limit theorem in certain settings. This enables asymptotically valid online inference, empowering real-time confidence interval construction. Furthermore, we propose an adaptive averaging scheme, inspired by optimal weights for linear models, which achieves optimal superior non-asymptotic bounds. Our theoretical advances and empirical validations redefine SGD’s capabilities, offering transformative insights for statistical learning and optimization.
APA
Wei, Z., Zhu, W. & Wu, W.B.. (2026). General Weighted Averaging in Stochastic Gradient Descent: CLT and Adaptive Optimality . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1198-1206 Available from https://proceedings.mlr.press/v300/wei26b.html.

Related Material