Probing the effect of data representation on transformer learning

Sarah McGuire Scullen, Henry Kvinge, Helen Jenne
Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026), PMLR 334(2):165-189, 2026.

Abstract

The choice of data representation can affect what a model learns, motivating the need for systematic methods to study this effect. Permutations and their statistics provide a controlled setting in which the underlying object is fixed and only its input representation varies. In this paper, we train small transformers on multiple representations of the same permutations and analyze them using linear probes, attention head ablations, and representational similarity measures. We find that convergence across input representations is task-dependent: models trained to compute the length of the longest increasing subsequence from different representations converge on similar internal computations, while models trained to compute other statistics show more varied behavior, including representation-specific algorithms. Beyond these specific findings, this work provides a controlled testbed for future interpretability research studying when models trained on different input representations learn the same mechanism.

Cite this Paper


BibTeX
@InProceedings{pmlr-v334-scullen26a, title = {Probing the effect of data representation on transformer learning}, author = {Scullen, Sarah McGuire and Kvinge, Henry and Jenne, Helen}, booktitle = {Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026)}, pages = {165--189}, year = {2026}, editor = {Berman, Eddie and Bernárdez, Guillermo and Chen, Samantha and Cloninger, Alex and Doster, Timothy and Emerson, Tegan and Grigsby, J. Elisenda and Kvinge, Henry and Lawrence, Hannah and Marrinan, Tim and Myers, Audun and Papillon, Mathilde and Tahmasebi, Behrooz and Telyatnikov, Lev and Walters, Robin and Weber, Melanie and Xie, YuQing and Yeats, Eric}, volume = {334}, number = {2}, series = {Proceedings of Machine Learning Research}, month = {18--20 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v334/main/assets/scullen26a/scullen26a.pdf}, url = {https://proceedings.mlr.press/v334/scullen26a.html}, abstract = {The choice of data representation can affect what a model learns, motivating the need for systematic methods to study this effect. Permutations and their statistics provide a controlled setting in which the underlying object is fixed and only its input representation varies. In this paper, we train small transformers on multiple representations of the same permutations and analyze them using linear probes, attention head ablations, and representational similarity measures. We find that convergence across input representations is task-dependent: models trained to compute the length of the longest increasing subsequence from different representations converge on similar internal computations, while models trained to compute other statistics show more varied behavior, including representation-specific algorithms. Beyond these specific findings, this work provides a controlled testbed for future interpretability research studying when models trained on different input representations learn the same mechanism.} }
Endnote
%0 Conference Paper %T Probing the effect of data representation on transformer learning %A Sarah McGuire Scullen %A Henry Kvinge %A Helen Jenne %B Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026) %C Proceedings of Machine Learning Research %D 2026 %E Eddie Berman %E Guillermo Bernárdez %E Samantha Chen %E Alex Cloninger %E Timothy Doster %E Tegan Emerson %E J. Elisenda Grigsby %E Henry Kvinge %E Hannah Lawrence %E Tim Marrinan %E Audun Myers %E Mathilde Papillon %E Behrooz Tahmasebi %E Lev Telyatnikov %E Robin Walters %E Melanie Weber %E YuQing Xie %E Eric Yeats %F pmlr-v334-scullen26a %I PMLR %P 165--189 %U https://proceedings.mlr.press/v334/scullen26a.html %V 334 %N 2 %X The choice of data representation can affect what a model learns, motivating the need for systematic methods to study this effect. Permutations and their statistics provide a controlled setting in which the underlying object is fixed and only its input representation varies. In this paper, we train small transformers on multiple representations of the same permutations and analyze them using linear probes, attention head ablations, and representational similarity measures. We find that convergence across input representations is task-dependent: models trained to compute the length of the longest increasing subsequence from different representations converge on similar internal computations, while models trained to compute other statistics show more varied behavior, including representation-specific algorithms. Beyond these specific findings, this work provides a controlled testbed for future interpretability research studying when models trained on different input representations learn the same mechanism.
APA
Scullen, S.M., Kvinge, H. & Jenne, H.. (2026). Probing the effect of data representation on transformer learning. Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026), in Proceedings of Machine Learning Research 334(2):165-189 Available from https://proceedings.mlr.press/v334/scullen26a.html.

Related Material