When Preference Labels Fall Short: Aligning Diffusion Models from Real Data

Weiyan Chen, Weijian Deng, Yao Xiao, Weijie Tu, Ziyi Dong, Ibrahim Radwan, Liang Lin, Pengxu Wei
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:13963-13979, 2026.

Abstract

Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existing approaches rely on preference pairs constructed from model-generated images. Such supervision is inherently relative and can be ambiguous when both samples exhibit artifacts or limited visual quality, making it difficult to infer what constitutes a truly desirable output. In this work, we investigate whether real data can serve as an alternative source of supervision for preference alignment. We adopt a data-centric perspective and study a curation strategy that treats real images as reference points and constructs preference signals by contrasting them with generated or perturbed samples, without requiring manually annotated preference pairs. Through empirical analysis, we show that real-data-based supervision provides effective guidance for aligning diffusion models and achieves performance comparable to existing preference-based methods. Our results suggest that real data offers a practical and complementary source of supervision for preference alignment and highlight directions of label-efficient alignment strategies. Code and models are available at https://cwyxx.github.io/RealAlign.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26s, title = {When Preference Labels Fall Short: Aligning Diffusion Models from Real Data}, author = {Chen, Weiyan and Deng, Weijian and Xiao, Yao and Tu, Weijie and Dong, Ziyi and Radwan, Ibrahim and Lin, Liang and Wei, Pengxu}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {13963--13979}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26s/chen26s.pdf}, url = {https://proceedings.mlr.press/v306/chen26s.html}, abstract = {Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existing approaches rely on preference pairs constructed from model-generated images. Such supervision is inherently relative and can be ambiguous when both samples exhibit artifacts or limited visual quality, making it difficult to infer what constitutes a truly desirable output. In this work, we investigate whether real data can serve as an alternative source of supervision for preference alignment. We adopt a data-centric perspective and study a curation strategy that treats real images as reference points and constructs preference signals by contrasting them with generated or perturbed samples, without requiring manually annotated preference pairs. Through empirical analysis, we show that real-data-based supervision provides effective guidance for aligning diffusion models and achieves performance comparable to existing preference-based methods. Our results suggest that real data offers a practical and complementary source of supervision for preference alignment and highlight directions of label-efficient alignment strategies. Code and models are available at https://cwyxx.github.io/RealAlign.} }
Endnote
%0 Conference Paper %T When Preference Labels Fall Short: Aligning Diffusion Models from Real Data %A Weiyan Chen %A Weijian Deng %A Yao Xiao %A Weijie Tu %A Ziyi Dong %A Ibrahim Radwan %A Liang Lin %A Pengxu Wei %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26s %I PMLR %P 13963--13979 %U https://proceedings.mlr.press/v306/chen26s.html %V 306 %X Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existing approaches rely on preference pairs constructed from model-generated images. Such supervision is inherently relative and can be ambiguous when both samples exhibit artifacts or limited visual quality, making it difficult to infer what constitutes a truly desirable output. In this work, we investigate whether real data can serve as an alternative source of supervision for preference alignment. We adopt a data-centric perspective and study a curation strategy that treats real images as reference points and constructs preference signals by contrasting them with generated or perturbed samples, without requiring manually annotated preference pairs. Through empirical analysis, we show that real-data-based supervision provides effective guidance for aligning diffusion models and achieves performance comparable to existing preference-based methods. Our results suggest that real data offers a practical and complementary source of supervision for preference alignment and highlight directions of label-efficient alignment strategies. Code and models are available at https://cwyxx.github.io/RealAlign.
APA
Chen, W., Deng, W., Xiao, Y., Tu, W., Dong, Z., Radwan, I., Lin, L. & Wei, P.. (2026). When Preference Labels Fall Short: Aligning Diffusion Models from Real Data. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:13963-13979 Available from https://proceedings.mlr.press/v306/chen26s.html.

Related Material