MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

Nicolas Dufour, Lucas Degeorge, Arijit Ghosh, Vicky Kalogeiton, David Picard
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27152-27200, 2026.

Abstract

The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-dufour26a, title = {{MIRO}: {M}ult{I}-Reward c{O}nditioned pretraining improves {T}2{I} quality and efficiency}, author = {Dufour, Nicolas and Degeorge, Lucas and Ghosh, Arijit and Kalogeiton, Vicky and Picard, David}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27152--27200}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/dufour26a/dufour26a.pdf}, url = {https://proceedings.mlr.press/v306/dufour26a.html}, abstract = {The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).} }
Endnote
%0 Conference Paper %T MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency %A Nicolas Dufour %A Lucas Degeorge %A Arijit Ghosh %A Vicky Kalogeiton %A David Picard %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-dufour26a %I PMLR %P 27152--27200 %U https://proceedings.mlr.press/v306/dufour26a.html %V 306 %X The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).
APA
Dufour, N., Degeorge, L., Ghosh, A., Kalogeiton, V. & Picard, D.. (2026). MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27152-27200 Available from https://proceedings.mlr.press/v306/dufour26a.html.

Related Material