Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization

Ryusei Yamada, Naoki Sato, Hideaki Iiduka
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7670-7700, 2026.

Abstract

Stochastic Gradient Descent ({SGD}) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla {SGD}, particularly with momentum, perform in the presence of heavy-tailed noise? In this paper, we refine existing convergence results for vanilla {SGD} and, more importantly, provide the first comprehensive convergence analysis of vanilla {SGD} with momentum for strongly convex, convex, and nonconvex objectives, without employing any gradient control mechanisms. Our results demonstrate that the obtained convergence rates are inferior to the optimal rates achieved by clipped or normalized variants of {SGD}, thereby revealing inherent limitations of vanilla methods under heavy-tailed noise. The theoretical findings are supported by experiments on synthetic functions.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-yamada26a, title = {Vanilla {SGD} with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization}, author = {Yamada, Ryusei and Sato, Naoki and Iiduka, Hideaki}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7670--7700}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/yamada26a/yamada26a.pdf}, url = {https://proceedings.mlr.press/v337/yamada26a.html}, abstract = {Stochastic Gradient Descent ({SGD}) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla {SGD}, particularly with momentum, perform in the presence of heavy-tailed noise? In this paper, we refine existing convergence results for vanilla {SGD} and, more importantly, provide the first comprehensive convergence analysis of vanilla {SGD} with momentum for strongly convex, convex, and nonconvex objectives, without employing any gradient control mechanisms. Our results demonstrate that the obtained convergence rates are inferior to the optimal rates achieved by clipped or normalized variants of {SGD}, thereby revealing inherent limitations of vanilla methods under heavy-tailed noise. The theoretical findings are supported by experiments on synthetic functions.} }
Endnote
%0 Conference Paper %T Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization %A Ryusei Yamada %A Naoki Sato %A Hideaki Iiduka %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-yamada26a %I PMLR %P 7670--7700 %U https://proceedings.mlr.press/v337/yamada26a.html %V 337 %X Stochastic Gradient Descent ({SGD}) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla {SGD}, particularly with momentum, perform in the presence of heavy-tailed noise? In this paper, we refine existing convergence results for vanilla {SGD} and, more importantly, provide the first comprehensive convergence analysis of vanilla {SGD} with momentum for strongly convex, convex, and nonconvex objectives, without employing any gradient control mechanisms. Our results demonstrate that the obtained convergence rates are inferior to the optimal rates achieved by clipped or normalized variants of {SGD}, thereby revealing inherent limitations of vanilla methods under heavy-tailed noise. The theoretical findings are supported by experiments on synthetic functions.
APA
Yamada, R., Sato, N. & Iiduka, H.. (2026). Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7670-7700 Available from https://proceedings.mlr.press/v337/yamada26a.html.

Related Material