Retrieval-Aware Distillation for Transformer-SSM Hybrids

Aviv Bick, Eric P. Xing, Albert Gu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:8235-8247, 2026.

Abstract

State-space models (SSMs) offer efficient sequence modeling but lag behind Transformers on benchmarks that require in-context retrieval. Prior work links this gap to a small set of attention heads, termed Gather-and-Aggregate (G&A), which SSMs struggle to reproduce. We propose retrieval-aware distillation, which converts a pretrained Transformer into a hybrid student by preserving only these retrieval-critical heads and distilling the rest into recurrent heads. We identify the essential heads via ablation on a synthetic retrieval task, producing a hybrid with sparse, non-uniform attention placement. We show that preserving just 2% of attention heads recovers over 95% of teacher performance on retrieval-heavy tasks (10 heads in a 1B model), requiring far fewer heads than hybrids that retain at least 25%. We further find that large recurrent states often compensate for missing retrieval: once retrieval is handled by these heads, the SSM backbone can be simplified with limited loss, even with an 8$\times$ reduction in state dimension. By reducing both the attention cache and the SSM state, the resulting hybrid is 5–6$\times$ more memory-efficient than comparable hybrids, closing the Transformer–SSM gap at a fraction of the memory cost.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bick26a, title = {Retrieval-Aware Distillation for Transformer-{SSM} Hybrids}, author = {Bick, Aviv and Xing, Eric P. and Gu, Albert}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {8235--8247}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bick26a/bick26a.pdf}, url = {https://proceedings.mlr.press/v306/bick26a.html}, abstract = {State-space models (SSMs) offer efficient sequence modeling but lag behind Transformers on benchmarks that require in-context retrieval. Prior work links this gap to a small set of attention heads, termed Gather-and-Aggregate (G&A), which SSMs struggle to reproduce. We propose retrieval-aware distillation, which converts a pretrained Transformer into a hybrid student by preserving only these retrieval-critical heads and distilling the rest into recurrent heads. We identify the essential heads via ablation on a synthetic retrieval task, producing a hybrid with sparse, non-uniform attention placement. We show that preserving just 2% of attention heads recovers over 95% of teacher performance on retrieval-heavy tasks (10 heads in a 1B model), requiring far fewer heads than hybrids that retain at least 25%. We further find that large recurrent states often compensate for missing retrieval: once retrieval is handled by these heads, the SSM backbone can be simplified with limited loss, even with an 8$\times$ reduction in state dimension. By reducing both the attention cache and the SSM state, the resulting hybrid is 5–6$\times$ more memory-efficient than comparable hybrids, closing the Transformer–SSM gap at a fraction of the memory cost.} }
Endnote
%0 Conference Paper %T Retrieval-Aware Distillation for Transformer-SSM Hybrids %A Aviv Bick %A Eric P. Xing %A Albert Gu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bick26a %I PMLR %P 8235--8247 %U https://proceedings.mlr.press/v306/bick26a.html %V 306 %X State-space models (SSMs) offer efficient sequence modeling but lag behind Transformers on benchmarks that require in-context retrieval. Prior work links this gap to a small set of attention heads, termed Gather-and-Aggregate (G&A), which SSMs struggle to reproduce. We propose retrieval-aware distillation, which converts a pretrained Transformer into a hybrid student by preserving only these retrieval-critical heads and distilling the rest into recurrent heads. We identify the essential heads via ablation on a synthetic retrieval task, producing a hybrid with sparse, non-uniform attention placement. We show that preserving just 2% of attention heads recovers over 95% of teacher performance on retrieval-heavy tasks (10 heads in a 1B model), requiring far fewer heads than hybrids that retain at least 25%. We further find that large recurrent states often compensate for missing retrieval: once retrieval is handled by these heads, the SSM backbone can be simplified with limited loss, even with an 8$\times$ reduction in state dimension. By reducing both the attention cache and the SSM state, the resulting hybrid is 5–6$\times$ more memory-efficient than comparable hybrids, closing the Transformer–SSM gap at a fraction of the memory cost.
APA
Bick, A., Xing, E.P. & Gu, A.. (2026). Retrieval-Aware Distillation for Transformer-SSM Hybrids. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:8235-8247 Available from https://proceedings.mlr.press/v306/bick26a.html.

Related Material