Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

Hwiyeong Lee, Ingyu Bang, Uiji Hwang, Hyelim Lim, Taeuk Kim
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:63392-63429, 2026.

Abstract

While sparse autoencoders provide features more interpretable than individual neurons, reliably characterizing them remains challenging. We propose Query Lens, which extends Logit Lens to enable more comprehensive and faithful interpretations of sparse features. By jointly considering encoder-side key features and decoder-side value features, we identify both the inputs that activate a feature and the outputs it promotes. We also account for indirect, module-mediated effects that arise when the feature is processed by downstream modules, going beyond the direct effect captured by Logit Lens. In experiments, we find that Query Lens yields coherent token signatures for features that remain uninterpretable under Logit Lens. Finally, we propose the Subspace Channel Hypothesis, suggesting that downstream modules read features through layer-specific subspaces.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-lee26a, title = {Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects}, author = {Lee, Hwiyeong and Bang, Ingyu and Hwang, Uiji and Lim, Hyelim and Kim, Taeuk}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {63392--63429}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/lee26a/lee26a.pdf}, url = {https://proceedings.mlr.press/v306/lee26a.html}, abstract = {While sparse autoencoders provide features more interpretable than individual neurons, reliably characterizing them remains challenging. We propose Query Lens, which extends Logit Lens to enable more comprehensive and faithful interpretations of sparse features. By jointly considering encoder-side key features and decoder-side value features, we identify both the inputs that activate a feature and the outputs it promotes. We also account for indirect, module-mediated effects that arise when the feature is processed by downstream modules, going beyond the direct effect captured by Logit Lens. In experiments, we find that Query Lens yields coherent token signatures for features that remain uninterpretable under Logit Lens. Finally, we propose the Subspace Channel Hypothesis, suggesting that downstream modules read features through layer-specific subspaces.} }
Endnote
%0 Conference Paper %T Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects %A Hwiyeong Lee %A Ingyu Bang %A Uiji Hwang %A Hyelim Lim %A Taeuk Kim %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-lee26a %I PMLR %P 63392--63429 %U https://proceedings.mlr.press/v306/lee26a.html %V 306 %X While sparse autoencoders provide features more interpretable than individual neurons, reliably characterizing them remains challenging. We propose Query Lens, which extends Logit Lens to enable more comprehensive and faithful interpretations of sparse features. By jointly considering encoder-side key features and decoder-side value features, we identify both the inputs that activate a feature and the outputs it promotes. We also account for indirect, module-mediated effects that arise when the feature is processed by downstream modules, going beyond the direct effect captured by Logit Lens. In experiments, we find that Query Lens yields coherent token signatures for features that remain uninterpretable under Logit Lens. Finally, we propose the Subspace Channel Hypothesis, suggesting that downstream modules read features through layer-specific subspaces.
APA
Lee, H., Bang, I., Hwang, U., Lim, H. & Kim, T.. (2026). Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:63392-63429 Available from https://proceedings.mlr.press/v306/lee26a.html.

Related Material