Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics

Raphael Bernas, Fanny Jourdan, Antonin Poché, Celine Hudelot
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:7741-7765, 2026.

Abstract

Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inherent anisotropy phenomenon in these models, presenting a significant challenge to their geometric interpretation. Previous theoretical studies on this phenomenon are rarely based on the underlying representation geometry. In this paper, we extend them by providing such theoretical arguments assessing the problematic nature of this phenomenon. Furthermore, to observe geometric internal model dynamics, we apply mechanistic interpretability (MI) techniques during the model’s training checkpoints rather than post-hoc, as it is commonly done in the literature. By analyzing multiple models and their checkpoints -including EuroBERT, the Pythia suite, and SmolLM2- we investigate the structure of embedding representations and their correlation with the on manifold entropy of their underlying distribution.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bernas26a, title = {Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics}, author = {Bernas, Raphael and Jourdan, Fanny and Poch\'{e}, Antonin and Hudelot, Celine}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {7741--7765}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bernas26a/bernas26a.pdf}, url = {https://proceedings.mlr.press/v306/bernas26a.html}, abstract = {Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inherent anisotropy phenomenon in these models, presenting a significant challenge to their geometric interpretation. Previous theoretical studies on this phenomenon are rarely based on the underlying representation geometry. In this paper, we extend them by providing such theoretical arguments assessing the problematic nature of this phenomenon. Furthermore, to observe geometric internal model dynamics, we apply mechanistic interpretability (MI) techniques during the model’s training checkpoints rather than post-hoc, as it is commonly done in the literature. By analyzing multiple models and their checkpoints -including EuroBERT, the Pythia suite, and SmolLM2- we investigate the structure of embedding representations and their correlation with the on manifold entropy of their underlying distribution.} }
Endnote
%0 Conference Paper %T Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics %A Raphael Bernas %A Fanny Jourdan %A Antonin Poché %A Celine Hudelot %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bernas26a %I PMLR %P 7741--7765 %U https://proceedings.mlr.press/v306/bernas26a.html %V 306 %X Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inherent anisotropy phenomenon in these models, presenting a significant challenge to their geometric interpretation. Previous theoretical studies on this phenomenon are rarely based on the underlying representation geometry. In this paper, we extend them by providing such theoretical arguments assessing the problematic nature of this phenomenon. Furthermore, to observe geometric internal model dynamics, we apply mechanistic interpretability (MI) techniques during the model’s training checkpoints rather than post-hoc, as it is commonly done in the literature. By analyzing multiple models and their checkpoints -including EuroBERT, the Pythia suite, and SmolLM2- we investigate the structure of embedding representations and their correlation with the on manifold entropy of their underlying distribution.
APA
Bernas, R., Jourdan, F., Poché, A. & Hudelot, C.. (2026). Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:7741-7765 Available from https://proceedings.mlr.press/v306/bernas26a.html.

Related Material