EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache Compression

Wenhao Gao, Haoran Cao, Yueyan Li, Yonggao Xiao, Caixia Yuan, Xiaojie Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:32994-33018, 2026.

Abstract

The prohibitive memory footprint of the Key-Value (KV) cache imposes a critical bottleneck for efficient long-context LLM serving. Current compression techniques typically rely on static or uniform budget allocation, overlooking the significant heterogeneity in information density across attention heads. To address this, we introduce EntroKV, an entropy-driven dynamic budget allocation framework. Our method enables dynamic and rational allocation across layers, attention heads, and different tasks. We demonstrate that attention entropy serves as a robust proxy for compression sensitivity: heads with high entropy require larger retention budgets, whereas low-entropy heads can be aggressively compressed without accuracy degradation. Functioning as a lightweight, plug-and-play module, EntroKV optimizes budget scheduling in real-time and is compatible with diverse compression operators. Extensive experiments demonstrate that EntroKV consistently outperforms baselines, retaining $\sim$98% of full-cache performance at a 30% budget ratio with negligible computational overhead. Our code is available at https://anonymous.4open.science/r/EntroKV-D0C8/.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gao26a, title = {{E}ntro{KV}: Entropy-Guided Dynamic Budget Allocation for {KV}-Cache Compression}, author = {Gao, Wenhao and Cao, Haoran and Li, Yueyan and Xiao, Yonggao and Yuan, Caixia and Wang, Xiaojie}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {32994--33018}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gao26a/gao26a.pdf}, url = {https://proceedings.mlr.press/v306/gao26a.html}, abstract = {The prohibitive memory footprint of the Key-Value (KV) cache imposes a critical bottleneck for efficient long-context LLM serving. Current compression techniques typically rely on static or uniform budget allocation, overlooking the significant heterogeneity in information density across attention heads. To address this, we introduce EntroKV, an entropy-driven dynamic budget allocation framework. Our method enables dynamic and rational allocation across layers, attention heads, and different tasks. We demonstrate that attention entropy serves as a robust proxy for compression sensitivity: heads with high entropy require larger retention budgets, whereas low-entropy heads can be aggressively compressed without accuracy degradation. Functioning as a lightweight, plug-and-play module, EntroKV optimizes budget scheduling in real-time and is compatible with diverse compression operators. Extensive experiments demonstrate that EntroKV consistently outperforms baselines, retaining $\sim$98% of full-cache performance at a 30% budget ratio with negligible computational overhead. Our code is available at https://anonymous.4open.science/r/EntroKV-D0C8/.} }
Endnote
%0 Conference Paper %T EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache Compression %A Wenhao Gao %A Haoran Cao %A Yueyan Li %A Yonggao Xiao %A Caixia Yuan %A Xiaojie Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gao26a %I PMLR %P 32994--33018 %U https://proceedings.mlr.press/v306/gao26a.html %V 306 %X The prohibitive memory footprint of the Key-Value (KV) cache imposes a critical bottleneck for efficient long-context LLM serving. Current compression techniques typically rely on static or uniform budget allocation, overlooking the significant heterogeneity in information density across attention heads. To address this, we introduce EntroKV, an entropy-driven dynamic budget allocation framework. Our method enables dynamic and rational allocation across layers, attention heads, and different tasks. We demonstrate that attention entropy serves as a robust proxy for compression sensitivity: heads with high entropy require larger retention budgets, whereas low-entropy heads can be aggressively compressed without accuracy degradation. Functioning as a lightweight, plug-and-play module, EntroKV optimizes budget scheduling in real-time and is compatible with diverse compression operators. Extensive experiments demonstrate that EntroKV consistently outperforms baselines, retaining $\sim$98% of full-cache performance at a 30% budget ratio with negligible computational overhead. Our code is available at https://anonymous.4open.science/r/EntroKV-D0C8/.
APA
Gao, W., Cao, H., Li, Y., Xiao, Y., Yuan, C. & Wang, X.. (2026). EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache Compression. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:32994-33018 Available from https://proceedings.mlr.press/v306/gao26a.html.

Related Material