Preserving Compositionality for Robust Multi-Subject Personalization in Text-to-Image Generation

Sangho Lee, Eugene Baek, Suho Ryu, Dongsoo Shin, Joonseok Lee
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3429-3452, 2026.

Abstract

Personalized text-to-image generation faces a fundamental trade-off between learning new concepts and preserving existing knowledge, and balancing these two goals gets particularly more challenging with multiple subjects to be personalized. Existing methods have relieved this by providing additional layout constraints, but it substantially undermines flexibility of the generation process. To address this issue, we propose a teacher–student architecture equipped with explicit regularizers to mitigate this trade-off. Specifically, we regulate the internal representations and cross-attention maps not to significantly deviate from the original foundation model, balanced with the reconstruction objective to internalize new concepts. Through extensive experiments, we demonstrate that our method enables reliable layout-free multi-subject generation, achieving the state-of-the-art performance on both single- and multi-concept personalization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-lee26f, title = {Preserving Compositionality for Robust Multi-Subject Personalization in Text-to-Image Generation}, author = {Lee, Sangho and Baek, Eugene and Ryu, Suho and Shin, Dongsoo and Lee, Joonseok}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3429--3452}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/lee26f/lee26f.pdf}, url = {https://proceedings.mlr.press/v337/lee26f.html}, abstract = {Personalized text-to-image generation faces a fundamental trade-off between learning new concepts and preserving existing knowledge, and balancing these two goals gets particularly more challenging with multiple subjects to be personalized. Existing methods have relieved this by providing additional layout constraints, but it substantially undermines flexibility of the generation process. To address this issue, we propose a teacher–student architecture equipped with explicit regularizers to mitigate this trade-off. Specifically, we regulate the internal representations and cross-attention maps not to significantly deviate from the original foundation model, balanced with the reconstruction objective to internalize new concepts. Through extensive experiments, we demonstrate that our method enables reliable layout-free multi-subject generation, achieving the state-of-the-art performance on both single- and multi-concept personalization.} }
Endnote
%0 Conference Paper %T Preserving Compositionality for Robust Multi-Subject Personalization in Text-to-Image Generation %A Sangho Lee %A Eugene Baek %A Suho Ryu %A Dongsoo Shin %A Joonseok Lee %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-lee26f %I PMLR %P 3429--3452 %U https://proceedings.mlr.press/v337/lee26f.html %V 337 %X Personalized text-to-image generation faces a fundamental trade-off between learning new concepts and preserving existing knowledge, and balancing these two goals gets particularly more challenging with multiple subjects to be personalized. Existing methods have relieved this by providing additional layout constraints, but it substantially undermines flexibility of the generation process. To address this issue, we propose a teacher–student architecture equipped with explicit regularizers to mitigate this trade-off. Specifically, we regulate the internal representations and cross-attention maps not to significantly deviate from the original foundation model, balanced with the reconstruction objective to internalize new concepts. Through extensive experiments, we demonstrate that our method enables reliable layout-free multi-subject generation, achieving the state-of-the-art performance on both single- and multi-concept personalization.
APA
Lee, S., Baek, E., Ryu, S., Shin, D. & Lee, J.. (2026). Preserving Compositionality for Robust Multi-Subject Personalization in Text-to-Image Generation. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3429-3452 Available from https://proceedings.mlr.press/v337/lee26f.html.

Related Material