Efficient Mismatch-Tolerant Coding for Model-Driven Compression

Aviv Adler, Jennifer Tang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:627-644, 2026.

Abstract

A central insight in lossless data compression is the close connection between probabilistic nextsymbol prediction and efficient sequence compression, whereby predictive models can be combined with classical coding techniques to achieve strong compression performance. Applying this approach with powerful modern learned models, such as LLMs, has been shown to achieve markedly better compression than traditional techniques across a wide range of domains. However, significant practical challenges remain, including model non-determinism, in which a model produces different predictions on different machines despite identical parameters and inputs; such mismatches between the encoder and decoder can lead to complete decoding failure. Probability Matching Interval Coding (PMATIC) was recently introduced as a drop-in framework for mismatch-robust coding and shown to enable reliable compression and decompression in the presence of bounded prediction mismatch (Adler & Tang, 2026). In this work, we present a generalization of PMATIC that allows the incorporation of tight theoretical results into the design and more flexible parameter optimization, resulting in substantial improvements in compression efficiency and robustness.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-adler26b, title = {Efficient Mismatch-Tolerant Coding for Model-Driven Compression}, author = {Adler, Aviv and Tang, Jennifer}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {627--644}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/adler26b/adler26b.pdf}, url = {https://proceedings.mlr.press/v306/adler26b.html}, abstract = {A central insight in lossless data compression is the close connection between probabilistic nextsymbol prediction and efficient sequence compression, whereby predictive models can be combined with classical coding techniques to achieve strong compression performance. Applying this approach with powerful modern learned models, such as LLMs, has been shown to achieve markedly better compression than traditional techniques across a wide range of domains. However, significant practical challenges remain, including model non-determinism, in which a model produces different predictions on different machines despite identical parameters and inputs; such mismatches between the encoder and decoder can lead to complete decoding failure. Probability Matching Interval Coding (PMATIC) was recently introduced as a drop-in framework for mismatch-robust coding and shown to enable reliable compression and decompression in the presence of bounded prediction mismatch (Adler & Tang, 2026). In this work, we present a generalization of PMATIC that allows the incorporation of tight theoretical results into the design and more flexible parameter optimization, resulting in substantial improvements in compression efficiency and robustness.} }
Endnote
%0 Conference Paper %T Efficient Mismatch-Tolerant Coding for Model-Driven Compression %A Aviv Adler %A Jennifer Tang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-adler26b %I PMLR %P 627--644 %U https://proceedings.mlr.press/v306/adler26b.html %V 306 %X A central insight in lossless data compression is the close connection between probabilistic nextsymbol prediction and efficient sequence compression, whereby predictive models can be combined with classical coding techniques to achieve strong compression performance. Applying this approach with powerful modern learned models, such as LLMs, has been shown to achieve markedly better compression than traditional techniques across a wide range of domains. However, significant practical challenges remain, including model non-determinism, in which a model produces different predictions on different machines despite identical parameters and inputs; such mismatches between the encoder and decoder can lead to complete decoding failure. Probability Matching Interval Coding (PMATIC) was recently introduced as a drop-in framework for mismatch-robust coding and shown to enable reliable compression and decompression in the presence of bounded prediction mismatch (Adler & Tang, 2026). In this work, we present a generalization of PMATIC that allows the incorporation of tight theoretical results into the design and more flexible parameter optimization, resulting in substantial improvements in compression efficiency and robustness.
APA
Adler, A. & Tang, J.. (2026). Efficient Mismatch-Tolerant Coding for Model-Driven Compression. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:627-644 Available from https://proceedings.mlr.press/v306/adler26b.html.

Related Material