Complete Characterization of Gauge Symmetries in Transformer Architectures

Hong Wang, Kelly Wang
Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations, PMLR 282:622-663, 2026.

Abstract

Modern Transformers possess redundant parameter symmetries that leave their function unchanged. We establish the complete gauge group structure for the canonical Transformer family, which encompasses standard architectures including GPT-2, BERT, LLaMA, and Qwen. For canonical Transformers with standard multi-head attention, we prove global maximality: the gauge group equals exactly $G_{max} = ((GL(d_k))^h \times (GL(d_v))^h) \rtimes S_h$ on the generic stratum where projection matrices have full column rank and head-wise attention controllability holds. For architectures with rotary position embeddings (RoPE) or relative encodings, as used in LLaMA and Qwen, the gauge group becomes $G_{RoPE} = ((C_{RoPE})^h \times (GL(d_v))^h) \rtimes S_h$ where $C_{RoPE}$ is the commutant of the position-dependent rotations—typically reducing to $(GL(1,\mathbb{C}))^{d_k/2}$ for standard RoPE implementations. We prove maximality through three key results: characterizing the Lie algebra of infinitesimal symmetries as $\mathfrak{g}_{max} = \bigoplus_{i=1}^h \mathfrak{gl}(d_k) \oplus \bigoplus_{i=1}^h \mathfrak{gl}(d_v)$ for canonical models, establishing that attention weights must be preserved up to head permutation under gauge equivalence, and demonstrating that query–key and value–output transformations necessarily factorize independently. These gauge symmetries persist through LayerNorm and extend to complete architectures, with the full model gauge group being $G_{Model} = \prod_{l=1}^L G_{Layer}^{(l)}$ Our characterization reveals over 1.1 million redundant dimensions in a 110M parameter Transformer Base model. Experiments on pretrained GPT-2 models from 124M to 1.5B parameters confirm that valid gauge transformations preserve model outputs to machine precision, while invalid transformations produce large errors, empirically supporting maximality.

Cite this Paper


BibTeX
@InProceedings{pmlr-v282-wang26a, title = {Complete Characterization of Gauge Symmetries in Transformer Architectures}, author = {Wang, Hong and Wang, Kelly}, booktitle = {Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations}, pages = {622--663}, year = {2026}, editor = {Acosta, Francisco and Azeglio, Simone and Tolooshams, Bahareh and van de Geijn, Chase and Shewmake, Christian and Sanborn, Sophia and Miolane, Nina}, volume = {282}, series = {Proceedings of Machine Learning Research}, month = {14 Dec 2024--07 Dec 2025}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v282/main/assets/wang26a/wang26a.pdf}, url = {https://proceedings.mlr.press/v282/wang26a.html}, abstract = {Modern Transformers possess redundant parameter symmetries that leave their function unchanged. We establish the complete gauge group structure for the canonical Transformer family, which encompasses standard architectures including GPT-2, BERT, LLaMA, and Qwen. For canonical Transformers with standard multi-head attention, we prove global maximality: the gauge group equals exactly $G_{max} = ((GL(d_k))^h \times (GL(d_v))^h) \rtimes S_h$ on the generic stratum where projection matrices have full column rank and head-wise attention controllability holds. For architectures with rotary position embeddings (RoPE) or relative encodings, as used in LLaMA and Qwen, the gauge group becomes $G_{RoPE} = ((C_{RoPE})^h \times (GL(d_v))^h) \rtimes S_h$ where $C_{RoPE}$ is the commutant of the position-dependent rotations—typically reducing to $(GL(1,\mathbb{C}))^{d_k/2}$ for standard RoPE implementations. We prove maximality through three key results: characterizing the Lie algebra of infinitesimal symmetries as $\mathfrak{g}_{max} = \bigoplus_{i=1}^h \mathfrak{gl}(d_k) \oplus \bigoplus_{i=1}^h \mathfrak{gl}(d_v)$ for canonical models, establishing that attention weights must be preserved up to head permutation under gauge equivalence, and demonstrating that query–key and value–output transformations necessarily factorize independently. These gauge symmetries persist through LayerNorm and extend to complete architectures, with the full model gauge group being $G_{Model} = \prod_{l=1}^L G_{Layer}^{(l)}$ Our characterization reveals over 1.1 million redundant dimensions in a 110M parameter Transformer Base model. Experiments on pretrained GPT-2 models from 124M to 1.5B parameters confirm that valid gauge transformations preserve model outputs to machine precision, while invalid transformations produce large errors, empirically supporting maximality.} }
Endnote
%0 Conference Paper %T Complete Characterization of Gauge Symmetries in Transformer Architectures %A Hong Wang %A Kelly Wang %B Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations %C Proceedings of Machine Learning Research %D 2026 %E Francisco Acosta %E Simone Azeglio %E Bahareh Tolooshams %E Chase van de Geijn %E Christian Shewmake %E Sophia Sanborn %E Nina Miolane %F pmlr-v282-wang26a %I PMLR %P 622--663 %U https://proceedings.mlr.press/v282/wang26a.html %V 282 %X Modern Transformers possess redundant parameter symmetries that leave their function unchanged. We establish the complete gauge group structure for the canonical Transformer family, which encompasses standard architectures including GPT-2, BERT, LLaMA, and Qwen. For canonical Transformers with standard multi-head attention, we prove global maximality: the gauge group equals exactly $G_{max} = ((GL(d_k))^h \times (GL(d_v))^h) \rtimes S_h$ on the generic stratum where projection matrices have full column rank and head-wise attention controllability holds. For architectures with rotary position embeddings (RoPE) or relative encodings, as used in LLaMA and Qwen, the gauge group becomes $G_{RoPE} = ((C_{RoPE})^h \times (GL(d_v))^h) \rtimes S_h$ where $C_{RoPE}$ is the commutant of the position-dependent rotations—typically reducing to $(GL(1,\mathbb{C}))^{d_k/2}$ for standard RoPE implementations. We prove maximality through three key results: characterizing the Lie algebra of infinitesimal symmetries as $\mathfrak{g}_{max} = \bigoplus_{i=1}^h \mathfrak{gl}(d_k) \oplus \bigoplus_{i=1}^h \mathfrak{gl}(d_v)$ for canonical models, establishing that attention weights must be preserved up to head permutation under gauge equivalence, and demonstrating that query–key and value–output transformations necessarily factorize independently. These gauge symmetries persist through LayerNorm and extend to complete architectures, with the full model gauge group being $G_{Model} = \prod_{l=1}^L G_{Layer}^{(l)}$ Our characterization reveals over 1.1 million redundant dimensions in a 110M parameter Transformer Base model. Experiments on pretrained GPT-2 models from 124M to 1.5B parameters confirm that valid gauge transformations preserve model outputs to machine precision, while invalid transformations produce large errors, empirically supporting maximality.
APA
Wang, H. & Wang, K.. (2026). Complete Characterization of Gauge Symmetries in Transformer Architectures. Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations, in Proceedings of Machine Learning Research 282:622-663 Available from https://proceedings.mlr.press/v282/wang26a.html.

Related Material