ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection

Debajyoti Datta, Trishala Neeraj, Bibek Paudel, Vyom Sharma, Subhabrata Mukherjee
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:23070-23095, 2026.

Abstract

Long-context inference is constrained by KV-cache memory, which grows linearly with sequence length; KV-cache compression therefore hinges on reliably selecting which past tokens to retain. Most geometry-based eviction methods score keys by cosine similarity to a global centroid, but cosine is scale-invariant and can discard magnitude cues that distinguish semantically salient tokens. We propose ManifoldKV, a training-free scorer that ranks tokens by Euclidean distance to the key centroid, capturing both angular and radial deviations. On the RULER benchmark, ManifoldKV achieves 95.7% accuracy at 4K–16K contexts with 20% compression, matching the best geometric baseline overall while decisively outperforming it in two regimes where magnitude information is critical. First, on multi-key retrieval, ManifoldKV reduces directional collisions, achieving 92.4% vs KeyDiff’s 77.0% (+15.4 points) on 3-key NIAH at 50% compression. Second, to address dilution and performance collapse of global centroids at 64K context, we introduce WindowedManifoldKV, which restores accuracy to 84.3% at 25% compression, a 49-point recovery over global L2 and +3.2 points over KeyDiff. Beyond RULER, we validate on real-world benchmarks: on LongBench, ManifoldKV outperforms KeyDiff by +2.80 points on Qwen3-8B (winning 12 of 14 tasks) and +0.49 on Phi-4; on HELMET, ManifoldKV achieves +6.5 EM on RAG and WindowedManifoldKV reaches +42 points on multi-key recall at 131K; and on InfiniteBench at 100K+ context, WindowedManifoldKV wins by +7.16 on Phi-4. Cross-architecture evaluation across six models reveals that the optimal distance metric depends on key-norm geometry, providing the first systematic guidelines for metric selection in geometric KV cache compression. The method requires only 3 lines of code and works across diverse architectures without tuning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-datta26b, title = {{M}anifold{KV}: Training-Free {KV} Cache Compression via {E}uclidean Outlier Detection}, author = {Datta, Debajyoti and Neeraj, Trishala and Paudel, Bibek and Sharma, Vyom and Mukherjee, Subhabrata}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {23070--23095}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/datta26b/datta26b.pdf}, url = {https://proceedings.mlr.press/v306/datta26b.html}, abstract = {Long-context inference is constrained by KV-cache memory, which grows linearly with sequence length; KV-cache compression therefore hinges on reliably selecting which past tokens to retain. Most geometry-based eviction methods score keys by cosine similarity to a global centroid, but cosine is scale-invariant and can discard magnitude cues that distinguish semantically salient tokens. We propose ManifoldKV, a training-free scorer that ranks tokens by Euclidean distance to the key centroid, capturing both angular and radial deviations. On the RULER benchmark, ManifoldKV achieves 95.7% accuracy at 4K–16K contexts with 20% compression, matching the best geometric baseline overall while decisively outperforming it in two regimes where magnitude information is critical. First, on multi-key retrieval, ManifoldKV reduces directional collisions, achieving 92.4% vs KeyDiff’s 77.0% (+15.4 points) on 3-key NIAH at 50% compression. Second, to address dilution and performance collapse of global centroids at 64K context, we introduce WindowedManifoldKV, which restores accuracy to 84.3% at 25% compression, a 49-point recovery over global L2 and +3.2 points over KeyDiff. Beyond RULER, we validate on real-world benchmarks: on LongBench, ManifoldKV outperforms KeyDiff by +2.80 points on Qwen3-8B (winning 12 of 14 tasks) and +0.49 on Phi-4; on HELMET, ManifoldKV achieves +6.5 EM on RAG and WindowedManifoldKV reaches +42 points on multi-key recall at 131K; and on InfiniteBench at 100K+ context, WindowedManifoldKV wins by +7.16 on Phi-4. Cross-architecture evaluation across six models reveals that the optimal distance metric depends on key-norm geometry, providing the first systematic guidelines for metric selection in geometric KV cache compression. The method requires only 3 lines of code and works across diverse architectures without tuning.} }
Endnote
%0 Conference Paper %T ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection %A Debajyoti Datta %A Trishala Neeraj %A Bibek Paudel %A Vyom Sharma %A Subhabrata Mukherjee %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-datta26b %I PMLR %P 23070--23095 %U https://proceedings.mlr.press/v306/datta26b.html %V 306 %X Long-context inference is constrained by KV-cache memory, which grows linearly with sequence length; KV-cache compression therefore hinges on reliably selecting which past tokens to retain. Most geometry-based eviction methods score keys by cosine similarity to a global centroid, but cosine is scale-invariant and can discard magnitude cues that distinguish semantically salient tokens. We propose ManifoldKV, a training-free scorer that ranks tokens by Euclidean distance to the key centroid, capturing both angular and radial deviations. On the RULER benchmark, ManifoldKV achieves 95.7% accuracy at 4K–16K contexts with 20% compression, matching the best geometric baseline overall while decisively outperforming it in two regimes where magnitude information is critical. First, on multi-key retrieval, ManifoldKV reduces directional collisions, achieving 92.4% vs KeyDiff’s 77.0% (+15.4 points) on 3-key NIAH at 50% compression. Second, to address dilution and performance collapse of global centroids at 64K context, we introduce WindowedManifoldKV, which restores accuracy to 84.3% at 25% compression, a 49-point recovery over global L2 and +3.2 points over KeyDiff. Beyond RULER, we validate on real-world benchmarks: on LongBench, ManifoldKV outperforms KeyDiff by +2.80 points on Qwen3-8B (winning 12 of 14 tasks) and +0.49 on Phi-4; on HELMET, ManifoldKV achieves +6.5 EM on RAG and WindowedManifoldKV reaches +42 points on multi-key recall at 131K; and on InfiniteBench at 100K+ context, WindowedManifoldKV wins by +7.16 on Phi-4. Cross-architecture evaluation across six models reveals that the optimal distance metric depends on key-norm geometry, providing the first systematic guidelines for metric selection in geometric KV cache compression. The method requires only 3 lines of code and works across diverse architectures without tuning.
APA
Datta, D., Neeraj, T., Paudel, B., Sharma, V. & Mukherjee, S.. (2026). ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:23070-23095 Available from https://proceedings.mlr.press/v306/datta26b.html.

Related Material