How Learning Dynamics Drive Adversarially Robust Generalization?

Yuelin Xu, Xiao Zhang
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7562-7593, 2026.

Abstract

Despite being widely adopted as a canonical framework for learning robust models, adversarial training suffers from robust overfitting. Existing empirical and theoretical explorations fail to provide a satisfactory mechanistic interpretation of the phenomenon. By modeling adversarial training with momentum {SGD} as a discrete-time dynamical system, we propose a {PAC}-{Bayesian} analytical framework that proves time-resolved robust generalization bounds. Specifically, our framework tracks the closed-form evolution of the posterior mean and covariance under both stationary and non-stationary transient regimes, connecting the model’s robust generalization performance to learning rate, local loss geometry, and mini-batch stochastic gradients. By estimating the key quantities associated with the bound, we illustrate the underlying mechanism of robust overfitting. Our framework also shows how adversarial weight perturbation reduces robust generalization gaps by suppressing dominant loss-curvature modes, while suggesting that excessive penalization can be sub-optimal for optimization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-xu26b, title = {How Learning Dynamics Drive Adversarially Robust Generalization?}, author = {Xu, Yuelin and Zhang, Xiao}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7562--7593}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/xu26b/xu26b.pdf}, url = {https://proceedings.mlr.press/v337/xu26b.html}, abstract = {Despite being widely adopted as a canonical framework for learning robust models, adversarial training suffers from robust overfitting. Existing empirical and theoretical explorations fail to provide a satisfactory mechanistic interpretation of the phenomenon. By modeling adversarial training with momentum {SGD} as a discrete-time dynamical system, we propose a {PAC}-{Bayesian} analytical framework that proves time-resolved robust generalization bounds. Specifically, our framework tracks the closed-form evolution of the posterior mean and covariance under both stationary and non-stationary transient regimes, connecting the model’s robust generalization performance to learning rate, local loss geometry, and mini-batch stochastic gradients. By estimating the key quantities associated with the bound, we illustrate the underlying mechanism of robust overfitting. Our framework also shows how adversarial weight perturbation reduces robust generalization gaps by suppressing dominant loss-curvature modes, while suggesting that excessive penalization can be sub-optimal for optimization.} }
Endnote
%0 Conference Paper %T How Learning Dynamics Drive Adversarially Robust Generalization? %A Yuelin Xu %A Xiao Zhang %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-xu26b %I PMLR %P 7562--7593 %U https://proceedings.mlr.press/v337/xu26b.html %V 337 %X Despite being widely adopted as a canonical framework for learning robust models, adversarial training suffers from robust overfitting. Existing empirical and theoretical explorations fail to provide a satisfactory mechanistic interpretation of the phenomenon. By modeling adversarial training with momentum {SGD} as a discrete-time dynamical system, we propose a {PAC}-{Bayesian} analytical framework that proves time-resolved robust generalization bounds. Specifically, our framework tracks the closed-form evolution of the posterior mean and covariance under both stationary and non-stationary transient regimes, connecting the model’s robust generalization performance to learning rate, local loss geometry, and mini-batch stochastic gradients. By estimating the key quantities associated with the bound, we illustrate the underlying mechanism of robust overfitting. Our framework also shows how adversarial weight perturbation reduces robust generalization gaps by suppressing dominant loss-curvature modes, while suggesting that excessive penalization can be sub-optimal for optimization.
APA
Xu, Y. & Zhang, X.. (2026). How Learning Dynamics Drive Adversarially Robust Generalization?. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7562-7593 Available from https://proceedings.mlr.press/v337/xu26b.html.

Related Material