ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning

Ge Gao, Di Xiong, Zeke Xie, Jian Yang, Shuo Chen
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:33703-33723, 2026.

Abstract

The unification of generative details and discriminative semantics presents a structural paradox in diffusion-based representation learning. Early approaches decouple semantics from generation, inevitably compromising representational completeness (i.e., information split). While recent bridge-based methods achieve unification via a tightly coupled mapping, they suffer from information overload. This is because unconstrained reconstruction objectives incentivize the encoder to entangle high-frequency stochastic noise into the latent bottleneck. To solve this, we introduce asymmetric rectified contrastive diffusion autoencoder (ArcDAE), which rebuilds the diffusion bridge as a dynamic sifter. Through imposing a timestep-aware rectification constraint that orthogonalizes the semantic manifold from the stochastic noise space, ArcDAE compels the bottleneck to distill discriminative features while actively shedding high-frequency redundancy. Consequently, our approach eliminates the overload trap without reverting to decoupling. Extensive experiments validate the superiority of our FFHQ-trained ArcDAE, surpassing state-of-the-art methods by up to 6.4% in downstream semantics regression and 9.7% in reconstruction fidelity.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gao26ad, title = {{A}rc{DAE}: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning}, author = {Gao, Ge and Xiong, Di and Xie, Zeke and Yang, Jian and Chen, Shuo}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {33703--33723}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gao26ad/gao26ad.pdf}, url = {https://proceedings.mlr.press/v306/gao26ad.html}, abstract = {The unification of generative details and discriminative semantics presents a structural paradox in diffusion-based representation learning. Early approaches decouple semantics from generation, inevitably compromising representational completeness (i.e., information split). While recent bridge-based methods achieve unification via a tightly coupled mapping, they suffer from information overload. This is because unconstrained reconstruction objectives incentivize the encoder to entangle high-frequency stochastic noise into the latent bottleneck. To solve this, we introduce asymmetric rectified contrastive diffusion autoencoder (ArcDAE), which rebuilds the diffusion bridge as a dynamic sifter. Through imposing a timestep-aware rectification constraint that orthogonalizes the semantic manifold from the stochastic noise space, ArcDAE compels the bottleneck to distill discriminative features while actively shedding high-frequency redundancy. Consequently, our approach eliminates the overload trap without reverting to decoupling. Extensive experiments validate the superiority of our FFHQ-trained ArcDAE, surpassing state-of-the-art methods by up to 6.4% in downstream semantics regression and 9.7% in reconstruction fidelity.} }
Endnote
%0 Conference Paper %T ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning %A Ge Gao %A Di Xiong %A Zeke Xie %A Jian Yang %A Shuo Chen %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gao26ad %I PMLR %P 33703--33723 %U https://proceedings.mlr.press/v306/gao26ad.html %V 306 %X The unification of generative details and discriminative semantics presents a structural paradox in diffusion-based representation learning. Early approaches decouple semantics from generation, inevitably compromising representational completeness (i.e., information split). While recent bridge-based methods achieve unification via a tightly coupled mapping, they suffer from information overload. This is because unconstrained reconstruction objectives incentivize the encoder to entangle high-frequency stochastic noise into the latent bottleneck. To solve this, we introduce asymmetric rectified contrastive diffusion autoencoder (ArcDAE), which rebuilds the diffusion bridge as a dynamic sifter. Through imposing a timestep-aware rectification constraint that orthogonalizes the semantic manifold from the stochastic noise space, ArcDAE compels the bottleneck to distill discriminative features while actively shedding high-frequency redundancy. Consequently, our approach eliminates the overload trap without reverting to decoupling. Extensive experiments validate the superiority of our FFHQ-trained ArcDAE, surpassing state-of-the-art methods by up to 6.4% in downstream semantics regression and 9.7% in reconstruction fidelity.
APA
Gao, G., Xiong, D., Xie, Z., Yang, J. & Chen, S.. (2026). ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:33703-33723 Available from https://proceedings.mlr.press/v306/gao26ad.html.

Related Material