Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models

Zongyu Guo, Jiajun He, Zhaoyang Jia, Xiaoyi Zhang, Jiahao Li, Xiao Li, Bin Li, José Miguel Hernández-Lobato, Yan Lu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:37909-37939, 2026.

Abstract

Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixels, latents, or tokens) remain external to the model and cannot directly exploit this knowledge for compact storage or reuse. In this work, we introduce a new visual representation framework that encodes a signal as a function, which is parametrized by low-rank adaptations attached to a frozen visual generative model. Such implicit representations of visual signals, e.g., an 81-frame video, can further be hashed into a single compact vector, achieving strong perceptual video compression at extremely low bitrates. Beyond basic compression, the functional nature of this representation enables inference-time scaling and control, allowing additional refinement on the compression performance. More broadly, as the implicit representations directly act as a function of the generation process, this suggests a unified framework bridging visual compression and generation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26e, title = {Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models}, author = {Guo, Zongyu and He, Jiajun and Jia, Zhaoyang and Zhang, Xiaoyi and Li, Jiahao and Li, Xiao and Li, Bin and Hern\'{a}ndez-Lobato, Jos\'{e} Miguel and Lu, Yan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {37909--37939}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26e/guo26e.pdf}, url = {https://proceedings.mlr.press/v306/guo26e.html}, abstract = {Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixels, latents, or tokens) remain external to the model and cannot directly exploit this knowledge for compact storage or reuse. In this work, we introduce a new visual representation framework that encodes a signal as a function, which is parametrized by low-rank adaptations attached to a frozen visual generative model. Such implicit representations of visual signals, e.g., an 81-frame video, can further be hashed into a single compact vector, achieving strong perceptual video compression at extremely low bitrates. Beyond basic compression, the functional nature of this representation enables inference-time scaling and control, allowing additional refinement on the compression performance. More broadly, as the implicit representations directly act as a function of the generation process, this suggests a unified framework bridging visual compression and generation.} }
Endnote
%0 Conference Paper %T Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models %A Zongyu Guo %A Jiajun He %A Zhaoyang Jia %A Xiaoyi Zhang %A Jiahao Li %A Xiao Li %A Bin Li %A José Miguel Hernández-Lobato %A Yan Lu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26e %I PMLR %P 37909--37939 %U https://proceedings.mlr.press/v306/guo26e.html %V 306 %X Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixels, latents, or tokens) remain external to the model and cannot directly exploit this knowledge for compact storage or reuse. In this work, we introduce a new visual representation framework that encodes a signal as a function, which is parametrized by low-rank adaptations attached to a frozen visual generative model. Such implicit representations of visual signals, e.g., an 81-frame video, can further be hashed into a single compact vector, achieving strong perceptual video compression at extremely low bitrates. Beyond basic compression, the functional nature of this representation enables inference-time scaling and control, allowing additional refinement on the compression performance. More broadly, as the implicit representations directly act as a function of the generation process, this suggests a unified framework bridging visual compression and generation.
APA
Guo, Z., He, J., Jia, Z., Zhang, X., Li, J., Li, X., Li, B., Hernández-Lobato, J.M. & Lu, Y.. (2026). Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:37909-37939 Available from https://proceedings.mlr.press/v306/guo26e.html.

Related Material