Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalization

Vitória Barin-Pacela, Shruti Joshi, Isabela Camacho, Simon Lacoste-Julien, David Klindt
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:364-412, 2026.

Abstract

Foundational to interpreting pretrained representations of deep generative models, the linear representation hypothesis states that neural network activations encode high-level concepts as linear mixtures. However, linear representation does not imply linear accessibility of such concepts: under superposition, when the number of concepts exceeds the activation dimension, recovering the underlying latent factors requires sparse nonlinear inference, making methods such as linear probes insufficient. Sparse autoencoders ({SAEs}) perform nonlinear inference but amortize it into a fixed encoder, introducing a systematic amortization gap. We show this gap dominates all other error sources and persists as the number of training samples is increased, causing {SAEs} to fail under out-of-distribution ({OOD}) compositional shifts. In contrast, classical sparse coding with per-sample iterative inference leverages compressed sensing guarantees to recover latent factors robustly, maintaining near-zero gaps in the accuracy between in and out of distribution. Our results demonstrate that the recent {OOD} failures of {SAEs} can be attributed to amortization failures: per-sample inference at test time substantially improves {OOD} performance, even when using a dictionary learned by an {SAE}. This is observed along a spectrum of hybrid approaches that progressively undo amortization and recover {OOD} performance.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-barin-pacela26a, title = {Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalization}, author = {Barin-Pacela, Vit\'{o}ria and Joshi, Shruti and Camacho, Isabela and Lacoste-Julien, Simon and Klindt, David}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {364--412}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/barin-pacela26a/barin-pacela26a.pdf}, url = {https://proceedings.mlr.press/v337/barin-pacela26a.html}, abstract = {Foundational to interpreting pretrained representations of deep generative models, the linear representation hypothesis states that neural network activations encode high-level concepts as linear mixtures. However, linear representation does not imply linear accessibility of such concepts: under superposition, when the number of concepts exceeds the activation dimension, recovering the underlying latent factors requires sparse nonlinear inference, making methods such as linear probes insufficient. Sparse autoencoders ({SAEs}) perform nonlinear inference but amortize it into a fixed encoder, introducing a systematic amortization gap. We show this gap dominates all other error sources and persists as the number of training samples is increased, causing {SAEs} to fail under out-of-distribution ({OOD}) compositional shifts. In contrast, classical sparse coding with per-sample iterative inference leverages compressed sensing guarantees to recover latent factors robustly, maintaining near-zero gaps in the accuracy between in and out of distribution. Our results demonstrate that the recent {OOD} failures of {SAEs} can be attributed to amortization failures: per-sample inference at test time substantially improves {OOD} performance, even when using a dictionary learned by an {SAE}. This is observed along a spectrum of hybrid approaches that progressively undo amortization and recover {OOD} performance.} }
Endnote
%0 Conference Paper %T Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalization %A Vitória Barin-Pacela %A Shruti Joshi %A Isabela Camacho %A Simon Lacoste-Julien %A David Klindt %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-barin-pacela26a %I PMLR %P 364--412 %U https://proceedings.mlr.press/v337/barin-pacela26a.html %V 337 %X Foundational to interpreting pretrained representations of deep generative models, the linear representation hypothesis states that neural network activations encode high-level concepts as linear mixtures. However, linear representation does not imply linear accessibility of such concepts: under superposition, when the number of concepts exceeds the activation dimension, recovering the underlying latent factors requires sparse nonlinear inference, making methods such as linear probes insufficient. Sparse autoencoders ({SAEs}) perform nonlinear inference but amortize it into a fixed encoder, introducing a systematic amortization gap. We show this gap dominates all other error sources and persists as the number of training samples is increased, causing {SAEs} to fail under out-of-distribution ({OOD}) compositional shifts. In contrast, classical sparse coding with per-sample iterative inference leverages compressed sensing guarantees to recover latent factors robustly, maintaining near-zero gaps in the accuracy between in and out of distribution. Our results demonstrate that the recent {OOD} failures of {SAEs} can be attributed to amortization failures: per-sample inference at test time substantially improves {OOD} performance, even when using a dictionary learned by an {SAE}. This is observed along a spectrum of hybrid approaches that progressively undo amortization and recover {OOD} performance.
APA
Barin-Pacela, V., Joshi, S., Camacho, I., Lacoste-Julien, S. & Klindt, D.. (2026). Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalization. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:364-412 Available from https://proceedings.mlr.press/v337/barin-pacela26a.html.

Related Material