The Information Geometry of Local Generalization Dynamics

Emmanouil M. Athanasakos
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3844-3852, 2026.

Abstract

Information-theoretic bounds on generalization are foundational to learning theory, yet their static form offers limited insight into the dynamic, iterative nature of modern optimization. This gap is addressed herein by developing a local theory of generalization based on Euclidean Information Theory, where each update is modeled as a perturbation vector. The resulting analysis shows that the change in the generalization gap is bounded by the expected squared norm of this vector, a quantity interpreted as the local generalization cost. The proposed local bound is proven to be the first-order approximation of classic global bounds, revealing their underlying differential structure. Within this framework, it is further revealed that the optimal update direction is mathematically equivalent to the natural gradient, offering an information-geometric justification for natural gradient descent. Finally, the overall theory is validated through experiments showing that the derived bound closely tracks training dynamics.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-athanasakos26a, title = { The Information Geometry of Local Generalization Dynamics }, author = {Athanasakos, Emmanouil M.}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3844--3852}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/athanasakos26a/athanasakos26a.pdf}, url = {https://proceedings.mlr.press/v300/athanasakos26a.html}, abstract = { Information-theoretic bounds on generalization are foundational to learning theory, yet their static form offers limited insight into the dynamic, iterative nature of modern optimization. This gap is addressed herein by developing a local theory of generalization based on Euclidean Information Theory, where each update is modeled as a perturbation vector. The resulting analysis shows that the change in the generalization gap is bounded by the expected squared norm of this vector, a quantity interpreted as the local generalization cost. The proposed local bound is proven to be the first-order approximation of classic global bounds, revealing their underlying differential structure. Within this framework, it is further revealed that the optimal update direction is mathematically equivalent to the natural gradient, offering an information-geometric justification for natural gradient descent. Finally, the overall theory is validated through experiments showing that the derived bound closely tracks training dynamics. } }
Endnote
%0 Conference Paper %T The Information Geometry of Local Generalization Dynamics %A Emmanouil M. Athanasakos %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-athanasakos26a %I PMLR %P 3844--3852 %U https://proceedings.mlr.press/v300/athanasakos26a.html %V 300 %X Information-theoretic bounds on generalization are foundational to learning theory, yet their static form offers limited insight into the dynamic, iterative nature of modern optimization. This gap is addressed herein by developing a local theory of generalization based on Euclidean Information Theory, where each update is modeled as a perturbation vector. The resulting analysis shows that the change in the generalization gap is bounded by the expected squared norm of this vector, a quantity interpreted as the local generalization cost. The proposed local bound is proven to be the first-order approximation of classic global bounds, revealing their underlying differential structure. Within this framework, it is further revealed that the optimal update direction is mathematically equivalent to the natural gradient, offering an information-geometric justification for natural gradient descent. Finally, the overall theory is validated through experiments showing that the derived bound closely tracks training dynamics.
APA
Athanasakos, E.M.. (2026). The Information Geometry of Local Generalization Dynamics . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3844-3852 Available from https://proceedings.mlr.press/v300/athanasakos26a.html.

Related Material