Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

Harry Jake Cunningham, Nicola Muca Cirone
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:22160-22180, 2026.

Abstract

Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approach has significant limitations as it neglects the geometric properties of the value vectors being aggregated. To address this gap, we introduce Contribution Weights, a projection-based metric that quantifies a token’s influence by accounting for it’s attention weight, value magnitude, and directional alignment with the layer output. We demonstrate that contribution weights provide a more faithful measure of token importance, consistently outperforming attention-based metrics in identifying semantically critical tokens across different decoder-only models, tasks, and datasets. Further, our metric enables novel mechanistic analysis of attention sinks. While previous work characterized sinks as passive repositories for excess attention, we reveal they serve an active functional role, suppressing information through a convex relationship between sink rate and output norm, stabilizing representations by opposing the semantic drift of low-confidence tokens.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cunningham26a, title = {Contribution Weights: A Geometrical Analysis of Self-Attention Transformers}, author = {Cunningham, Harry Jake and Muca Cirone, Nicola}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {22160--22180}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cunningham26a/cunningham26a.pdf}, url = {https://proceedings.mlr.press/v306/cunningham26a.html}, abstract = {Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approach has significant limitations as it neglects the geometric properties of the value vectors being aggregated. To address this gap, we introduce Contribution Weights, a projection-based metric that quantifies a token’s influence by accounting for it’s attention weight, value magnitude, and directional alignment with the layer output. We demonstrate that contribution weights provide a more faithful measure of token importance, consistently outperforming attention-based metrics in identifying semantically critical tokens across different decoder-only models, tasks, and datasets. Further, our metric enables novel mechanistic analysis of attention sinks. While previous work characterized sinks as passive repositories for excess attention, we reveal they serve an active functional role, suppressing information through a convex relationship between sink rate and output norm, stabilizing representations by opposing the semantic drift of low-confidence tokens.} }
Endnote
%0 Conference Paper %T Contribution Weights: A Geometrical Analysis of Self-Attention Transformers %A Harry Jake Cunningham %A Nicola Muca Cirone %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cunningham26a %I PMLR %P 22160--22180 %U https://proceedings.mlr.press/v306/cunningham26a.html %V 306 %X Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approach has significant limitations as it neglects the geometric properties of the value vectors being aggregated. To address this gap, we introduce Contribution Weights, a projection-based metric that quantifies a token’s influence by accounting for it’s attention weight, value magnitude, and directional alignment with the layer output. We demonstrate that contribution weights provide a more faithful measure of token importance, consistently outperforming attention-based metrics in identifying semantically critical tokens across different decoder-only models, tasks, and datasets. Further, our metric enables novel mechanistic analysis of attention sinks. While previous work characterized sinks as passive repositories for excess attention, we reveal they serve an active functional role, suppressing information through a convex relationship between sink rate and output norm, stabilizing representations by opposing the semantic drift of low-confidence tokens.
APA
Cunningham, H.J. & Muca Cirone, N.. (2026). Contribution Weights: A Geometrical Analysis of Self-Attention Transformers. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:22160-22180 Available from https://proceedings.mlr.press/v306/cunningham26a.html.

Related Material