Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

Shilong Zhang, He Zhang, Zhifei Zhang, Chongjian Ge, Shuchen Xue, Shaoteng Liu, Mengwei Ren, Soo Ye Kim, Yuqian Zhou, Qing Liu, Daniil Pakhomov, Kai Zhang, Zhe Lin, Ping Luo
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:160756-160772, 2026.

Abstract

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder’s inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic–pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with $16\times$ spatial downsampling). This design allows the latent space to remain semantically rich while achieving state-of-the-art image reconstruction, and keeps it compact enough for accurate generation. Leveraging this representation, we design a unified text-to-image (T2I) and image editing model. Across diverse generation spaces, our approach achieves state-of-the-art reconstruction, faster convergence, and substantial gains in both T2I and editing tasks, demonstrating that representation encoders can be effectively adapted into robust generative components.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-zhang26iw, title = {Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing}, author = {Zhang, Shilong and Zhang, He and Zhang, Zhifei and Ge, Chongjian and Xue, Shuchen and Liu, Shaoteng and Ren, Mengwei and Kim, Soo Ye and Zhou, Yuqian and Liu, Qing and Pakhomov, Daniil and Zhang, Kai and Lin, Zhe and Luo, Ping}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {160756--160772}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/zhang26iw/zhang26iw.pdf}, url = {https://proceedings.mlr.press/v306/zhang26iw.html}, abstract = {Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder’s inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic–pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with $16\times$ spatial downsampling). This design allows the latent space to remain semantically rich while achieving state-of-the-art image reconstruction, and keeps it compact enough for accurate generation. Leveraging this representation, we design a unified text-to-image (T2I) and image editing model. Across diverse generation spaces, our approach achieves state-of-the-art reconstruction, faster convergence, and substantial gains in both T2I and editing tasks, demonstrating that representation encoders can be effectively adapted into robust generative components.} }
Endnote
%0 Conference Paper %T Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing %A Shilong Zhang %A He Zhang %A Zhifei Zhang %A Chongjian Ge %A Shuchen Xue %A Shaoteng Liu %A Mengwei Ren %A Soo Ye Kim %A Yuqian Zhou %A Qing Liu %A Daniil Pakhomov %A Kai Zhang %A Zhe Lin %A Ping Luo %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-zhang26iw %I PMLR %P 160756--160772 %U https://proceedings.mlr.press/v306/zhang26iw.html %V 306 %X Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder’s inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic–pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with $16\times$ spatial downsampling). This design allows the latent space to remain semantically rich while achieving state-of-the-art image reconstruction, and keeps it compact enough for accurate generation. Leveraging this representation, we design a unified text-to-image (T2I) and image editing model. Across diverse generation spaces, our approach achieves state-of-the-art reconstruction, faster convergence, and substantial gains in both T2I and editing tasks, demonstrating that representation encoders can be effectively adapted into robust generative components.
APA
Zhang, S., Zhang, H., Zhang, Z., Ge, C., Xue, S., Liu, S., Ren, M., Kim, S.Y., Zhou, Y., Liu, Q., Pakhomov, D., Zhang, K., Lin, Z. & Luo, P.. (2026). Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:160756-160772 Available from https://proceedings.mlr.press/v306/zhang26iw.html.

Related Material