xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Hung-Yueh Chiang, Yash Akhauri, Xilai Dai, Huiqiang Jiang, Yucheng Li, Luis Ceze, Kai-Chiang Wu, Mohamed S. Abdelfattah
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:12758-12778, 2026.

Abstract

Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent studies attempt to share KV-Cache across layers, but these approaches either require expensive pretraining or rely on per-token cross-layer cosine similarity that is often limited in practice. We show, via Centered Kernel Alignment (CKA), that the dominant singular vectors of KV-Cache are well aligned across layers. Motivated by this observation, we propose xKV, a post-training compression method that jointly factorizes grouped-layer KV-Cache into a shared low-rank subspace, substantially reducing KV-Cache memory. Across widely used LLMs, xKV achieves up to 8$\times$ KV-Cache compression while preserving accuracy on long-context tasks and in multi-turn settings. To further improve efficiency, we introduce Selective Reconstruction (SR) at decode time. Combined with SR, xKV achieves up to 4.23$\times$ end-to-end speedup over the full attention baseline, and surpasses notable baselines with 30% higher throughput under a similar accuracy level. Overall, xKV provides a plug-and-play approach to reduce both memory and latency for long-context LLM inference. Our code is publicly available at: https://github.com/abdelfattah-lab/xKV.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chang26d, title = {x{KV}: Cross-Layer {KV}-Cache Compression via Aligned Singular Vector Extraction}, author = {Chang, Chi-Chih and Lin, Wei-Cheng and Lin, Chien-Yu and Chiang, Hung-Yueh and Akhauri, Yash and Dai, Xilai and Jiang, Huiqiang and Li, Yucheng and Ceze, Luis and Wu, Kai-Chiang and Abdelfattah, Mohamed S.}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {12758--12778}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chang26d/chang26d.pdf}, url = {https://proceedings.mlr.press/v306/chang26d.html}, abstract = {Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent studies attempt to share KV-Cache across layers, but these approaches either require expensive pretraining or rely on per-token cross-layer cosine similarity that is often limited in practice. We show, via Centered Kernel Alignment (CKA), that the dominant singular vectors of KV-Cache are well aligned across layers. Motivated by this observation, we propose xKV, a post-training compression method that jointly factorizes grouped-layer KV-Cache into a shared low-rank subspace, substantially reducing KV-Cache memory. Across widely used LLMs, xKV achieves up to 8$\times$ KV-Cache compression while preserving accuracy on long-context tasks and in multi-turn settings. To further improve efficiency, we introduce Selective Reconstruction (SR) at decode time. Combined with SR, xKV achieves up to 4.23$\times$ end-to-end speedup over the full attention baseline, and surpasses notable baselines with 30% higher throughput under a similar accuracy level. Overall, xKV provides a plug-and-play approach to reduce both memory and latency for long-context LLM inference. Our code is publicly available at: https://github.com/abdelfattah-lab/xKV.} }
Endnote
%0 Conference Paper %T xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction %A Chi-Chih Chang %A Wei-Cheng Lin %A Chien-Yu Lin %A Hung-Yueh Chiang %A Yash Akhauri %A Xilai Dai %A Huiqiang Jiang %A Yucheng Li %A Luis Ceze %A Kai-Chiang Wu %A Mohamed S. Abdelfattah %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chang26d %I PMLR %P 12758--12778 %U https://proceedings.mlr.press/v306/chang26d.html %V 306 %X Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent studies attempt to share KV-Cache across layers, but these approaches either require expensive pretraining or rely on per-token cross-layer cosine similarity that is often limited in practice. We show, via Centered Kernel Alignment (CKA), that the dominant singular vectors of KV-Cache are well aligned across layers. Motivated by this observation, we propose xKV, a post-training compression method that jointly factorizes grouped-layer KV-Cache into a shared low-rank subspace, substantially reducing KV-Cache memory. Across widely used LLMs, xKV achieves up to 8$\times$ KV-Cache compression while preserving accuracy on long-context tasks and in multi-turn settings. To further improve efficiency, we introduce Selective Reconstruction (SR) at decode time. Combined with SR, xKV achieves up to 4.23$\times$ end-to-end speedup over the full attention baseline, and surpasses notable baselines with 30% higher throughput under a similar accuracy level. Overall, xKV provides a plug-and-play approach to reduce both memory and latency for long-context LLM inference. Our code is publicly available at: https://github.com/abdelfattah-lab/xKV.
APA
Chang, C., Lin, W., Lin, C., Chiang, H., Akhauri, Y., Dai, X., Jiang, H., Li, Y., Ceze, L., Wu, K. & Abdelfattah, M.S.. (2026). xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:12758-12778 Available from https://proceedings.mlr.press/v306/chang26d.html.

Related Material