[edit]
Preserving Compositionality for Robust Multi-Subject Personalization in Text-to-Image Generation
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3429-3452, 2026.
Abstract
Personalized text-to-image generation faces a fundamental trade-off between learning new concepts and preserving existing knowledge, and balancing these two goals gets particularly more challenging with multiple subjects to be personalized. Existing methods have relieved this by providing additional layout constraints, but it substantially undermines flexibility of the generation process. To address this issue, we propose a teacher–student architecture equipped with explicit regularizers to mitigate this trade-off. Specifically, we regulate the internal representations and cross-attention maps not to significantly deviate from the original foundation model, balanced with the reconstruction objective to internalize new concepts. Through extensive experiments, we demonstrate that our method enables reliable layout-free multi-subject generation, achieving the state-of-the-art performance on both single- and multi-concept personalization.