Same Benchmark, Same Subspace: Task-Selective Convergence in LLM Representations

JaeSeong Kim, Suan Lee
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3101-3116, 2026.

Abstract

Benchmark scores are the dominant lens through which progress in large language models is assessed, yet the implicit assumption that different benchmarks read out a shared internal structure has never been verified at the representation level. We introduce readout subspace analysis, a framework that extracts the low-dimensional subspace a benchmark uses to linearly discriminate correct answers from a model’s frozen hidden representations, and compares these subspaces across models and tasks via a coordinate-invariant Gram cosine metric. Analyzing 50 models spanning six architecture families and varied fine-tuning strategies (SFT, DPO, RLHF, model merging) on three benchmarks (MMLU, MedQA, AGIEval), we find that readout geometry is governed by the benchmark, not the model. Within the same benchmark, readout subspaces are strongly aligned even across architecturally unrelated models (Gram cosine up to 0.87 versus a random baseline of ${\sim}10^{-3}$), while subspaces across different benchmarks are nearly orthogonal. This convergence is graded: benchmarks sharing cognitive demands partially overlap in readout subspace, whereas those requiring qualitatively different reasoning occupy orthogonal regions. Furthermore, probe accuracy consistently meets or exceeds generation accuracy, indicating that benchmark scores reflect not only representational structure but also decoding efficiency. These findings recast benchmarks from passive measurement instruments to active structural constraints that selectively activate specific regions of representation space.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kim26g, title = {Same Benchmark, Same Subspace: Task-Selective Convergence in {LLM} Representations}, author = {Kim, JaeSeong and Lee, Suan}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3101--3116}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kim26g/kim26g.pdf}, url = {https://proceedings.mlr.press/v337/kim26g.html}, abstract = {Benchmark scores are the dominant lens through which progress in large language models is assessed, yet the implicit assumption that different benchmarks read out a shared internal structure has never been verified at the representation level. We introduce readout subspace analysis, a framework that extracts the low-dimensional subspace a benchmark uses to linearly discriminate correct answers from a model’s frozen hidden representations, and compares these subspaces across models and tasks via a coordinate-invariant Gram cosine metric. Analyzing 50 models spanning six architecture families and varied fine-tuning strategies (SFT, DPO, RLHF, model merging) on three benchmarks (MMLU, MedQA, AGIEval), we find that readout geometry is governed by the benchmark, not the model. Within the same benchmark, readout subspaces are strongly aligned even across architecturally unrelated models (Gram cosine up to 0.87 versus a random baseline of ${\sim}10^{-3}$), while subspaces across different benchmarks are nearly orthogonal. This convergence is graded: benchmarks sharing cognitive demands partially overlap in readout subspace, whereas those requiring qualitatively different reasoning occupy orthogonal regions. Furthermore, probe accuracy consistently meets or exceeds generation accuracy, indicating that benchmark scores reflect not only representational structure but also decoding efficiency. These findings recast benchmarks from passive measurement instruments to active structural constraints that selectively activate specific regions of representation space.} }
Endnote
%0 Conference Paper %T Same Benchmark, Same Subspace: Task-Selective Convergence in LLM Representations %A JaeSeong Kim %A Suan Lee %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kim26g %I PMLR %P 3101--3116 %U https://proceedings.mlr.press/v337/kim26g.html %V 337 %X Benchmark scores are the dominant lens through which progress in large language models is assessed, yet the implicit assumption that different benchmarks read out a shared internal structure has never been verified at the representation level. We introduce readout subspace analysis, a framework that extracts the low-dimensional subspace a benchmark uses to linearly discriminate correct answers from a model’s frozen hidden representations, and compares these subspaces across models and tasks via a coordinate-invariant Gram cosine metric. Analyzing 50 models spanning six architecture families and varied fine-tuning strategies (SFT, DPO, RLHF, model merging) on three benchmarks (MMLU, MedQA, AGIEval), we find that readout geometry is governed by the benchmark, not the model. Within the same benchmark, readout subspaces are strongly aligned even across architecturally unrelated models (Gram cosine up to 0.87 versus a random baseline of ${\sim}10^{-3}$), while subspaces across different benchmarks are nearly orthogonal. This convergence is graded: benchmarks sharing cognitive demands partially overlap in readout subspace, whereas those requiring qualitatively different reasoning occupy orthogonal regions. Furthermore, probe accuracy consistently meets or exceeds generation accuracy, indicating that benchmark scores reflect not only representational structure but also decoding efficiency. These findings recast benchmarks from passive measurement instruments to active structural constraints that selectively activate specific regions of representation space.
APA
Kim, J. & Lee, S.. (2026). Same Benchmark, Same Subspace: Task-Selective Convergence in LLM Representations. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3101-3116 Available from https://proceedings.mlr.press/v337/kim26g.html.

Related Material