The Convergence Behavior of Adam under Heavy-Tailed Noise

Yijiang Pang
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:5244-5265, 2026.

Abstract

We establish the first convergence guarantees for the plain vector-form \emph{{Adam}} optimizer under heavy-tailed stochastic noise. While several {Adam} variants are known to achieve optimal iteration complexity in bounded-variance nonconvex optimization, little is understood about their behavior when stochastic gradients admit only a bounded $p$-th central moment for some $p \in (1,2]$, a setting increasingly observed in modern deep learning. To address this gap, we generalize the recent online-to-nonconvex conversion framework to accommodate heavy-tailed martingale-difference noise. Building on this generalized framework, we develop a discounted regret analysis for {Adam}, without restrictive parameter coupling. Our results show that {Adam} converges to $(\rho,\epsilon)$-stationary points under heavy-tailed noise. However, it exhibits a suboptimal iteration complexity and $p$-dependent convergence, a suboptimality that persists even in the bounded-variance case ($p=2$). When the domain radius is known and used to control the online-learner output, a standard setup in related literature, the convergence rate improves to match the optimal complexity. These findings provide new theoretical insight into the robustness and limitations of {Adam} in heavy-tailed regimes.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-pang26a, title = {The Convergence Behavior of {Adam} under Heavy-Tailed Noise}, author = {Pang, Yijiang}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {5244--5265}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/pang26a/pang26a.pdf}, url = {https://proceedings.mlr.press/v337/pang26a.html}, abstract = {We establish the first convergence guarantees for the plain vector-form \emph{{Adam}} optimizer under heavy-tailed stochastic noise. While several {Adam} variants are known to achieve optimal iteration complexity in bounded-variance nonconvex optimization, little is understood about their behavior when stochastic gradients admit only a bounded $p$-th central moment for some $p \in (1,2]$, a setting increasingly observed in modern deep learning. To address this gap, we generalize the recent online-to-nonconvex conversion framework to accommodate heavy-tailed martingale-difference noise. Building on this generalized framework, we develop a discounted regret analysis for {Adam}, without restrictive parameter coupling. Our results show that {Adam} converges to $(\rho,\epsilon)$-stationary points under heavy-tailed noise. However, it exhibits a suboptimal iteration complexity and $p$-dependent convergence, a suboptimality that persists even in the bounded-variance case ($p=2$). When the domain radius is known and used to control the online-learner output, a standard setup in related literature, the convergence rate improves to match the optimal complexity. These findings provide new theoretical insight into the robustness and limitations of {Adam} in heavy-tailed regimes.} }
Endnote
%0 Conference Paper %T The Convergence Behavior of Adam under Heavy-Tailed Noise %A Yijiang Pang %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-pang26a %I PMLR %P 5244--5265 %U https://proceedings.mlr.press/v337/pang26a.html %V 337 %X We establish the first convergence guarantees for the plain vector-form \emph{{Adam}} optimizer under heavy-tailed stochastic noise. While several {Adam} variants are known to achieve optimal iteration complexity in bounded-variance nonconvex optimization, little is understood about their behavior when stochastic gradients admit only a bounded $p$-th central moment for some $p \in (1,2]$, a setting increasingly observed in modern deep learning. To address this gap, we generalize the recent online-to-nonconvex conversion framework to accommodate heavy-tailed martingale-difference noise. Building on this generalized framework, we develop a discounted regret analysis for {Adam}, without restrictive parameter coupling. Our results show that {Adam} converges to $(\rho,\epsilon)$-stationary points under heavy-tailed noise. However, it exhibits a suboptimal iteration complexity and $p$-dependent convergence, a suboptimality that persists even in the bounded-variance case ($p=2$). When the domain radius is known and used to control the online-learner output, a standard setup in related literature, the convergence rate improves to match the optimal complexity. These findings provide new theoretical insight into the robustness and limitations of {Adam} in heavy-tailed regimes.
APA
Pang, Y.. (2026). The Convergence Behavior of Adam under Heavy-Tailed Noise. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:5244-5265 Available from https://proceedings.mlr.press/v337/pang26a.html.

Related Material