[edit]
ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:33703-33723, 2026.
Abstract
The unification of generative details and discriminative semantics presents a structural paradox in diffusion-based representation learning. Early approaches decouple semantics from generation, inevitably compromising representational completeness (i.e., information split). While recent bridge-based methods achieve unification via a tightly coupled mapping, they suffer from information overload. This is because unconstrained reconstruction objectives incentivize the encoder to entangle high-frequency stochastic noise into the latent bottleneck. To solve this, we introduce asymmetric rectified contrastive diffusion autoencoder (ArcDAE), which rebuilds the diffusion bridge as a dynamic sifter. Through imposing a timestep-aware rectification constraint that orthogonalizes the semantic manifold from the stochastic noise space, ArcDAE compels the bottleneck to distill discriminative features while actively shedding high-frequency redundancy. Consequently, our approach eliminates the overload trap without reverting to decoupling. Extensive experiments validate the superiority of our FFHQ-trained ArcDAE, surpassing state-of-the-art methods by up to 6.4% in downstream semantics regression and 9.7% in reconstruction fidelity.