StructMAR: Structure-Aware Masked Autoregression for Explicit Layout Alignment in Text-to-Image Generation

Gang Cao, Junying Zhang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:11720-11742, 2026.

Abstract

Although text-to-image generation has achieved significant progress, strict instance-level layout alignment remains a challenge for many applications. Masked Autoregressive (MAR) models on continuous latents are both efficient and high-fidelity, yet the standard practice of flattening 2D latents into 1D sequences weakens spatial topology, limiting precise controllability. To address this, we propose StructMAR, a structure-aware masked autoregressive framework that transforms layout alignment from a soft correlation into an explicit structural alignment. By integrating 2D Rotary Positional Embeddings with a Layout-Guided Attention Bias, StructMAR explicitly biases latent tokens toward their corresponding layout instances during attention computation. We further use Group Relative Policy Optimization (GRPO) as a final-stage policy refinement to reduce the mismatch between the MAR training objective and detector-based layout evaluation metrics. Evaluated on the COCO-Position and COCO-MIG benchmarks, StructMAR achieves state-of-the-art performance, reaching 57.2 AP and 79.4 mIoU on the former, and 61.7 ISR and 56.9 mIoU on the latter. These results, coupled with a 4.05$\times$ inference speedup, underscore the efficacy of explicit structural inductive biases in controllable autoregressive generation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cao26ab, title = {{S}truct{MAR}: Structure-Aware Masked Autoregression for Explicit Layout Alignment in Text-to-Image Generation}, author = {Cao, Gang and Zhang, Junying}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {11720--11742}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cao26ab/cao26ab.pdf}, url = {https://proceedings.mlr.press/v306/cao26ab.html}, abstract = {Although text-to-image generation has achieved significant progress, strict instance-level layout alignment remains a challenge for many applications. Masked Autoregressive (MAR) models on continuous latents are both efficient and high-fidelity, yet the standard practice of flattening 2D latents into 1D sequences weakens spatial topology, limiting precise controllability. To address this, we propose StructMAR, a structure-aware masked autoregressive framework that transforms layout alignment from a soft correlation into an explicit structural alignment. By integrating 2D Rotary Positional Embeddings with a Layout-Guided Attention Bias, StructMAR explicitly biases latent tokens toward their corresponding layout instances during attention computation. We further use Group Relative Policy Optimization (GRPO) as a final-stage policy refinement to reduce the mismatch between the MAR training objective and detector-based layout evaluation metrics. Evaluated on the COCO-Position and COCO-MIG benchmarks, StructMAR achieves state-of-the-art performance, reaching 57.2 AP and 79.4 mIoU on the former, and 61.7 ISR and 56.9 mIoU on the latter. These results, coupled with a 4.05$\times$ inference speedup, underscore the efficacy of explicit structural inductive biases in controllable autoregressive generation.} }
Endnote
%0 Conference Paper %T StructMAR: Structure-Aware Masked Autoregression for Explicit Layout Alignment in Text-to-Image Generation %A Gang Cao %A Junying Zhang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cao26ab %I PMLR %P 11720--11742 %U https://proceedings.mlr.press/v306/cao26ab.html %V 306 %X Although text-to-image generation has achieved significant progress, strict instance-level layout alignment remains a challenge for many applications. Masked Autoregressive (MAR) models on continuous latents are both efficient and high-fidelity, yet the standard practice of flattening 2D latents into 1D sequences weakens spatial topology, limiting precise controllability. To address this, we propose StructMAR, a structure-aware masked autoregressive framework that transforms layout alignment from a soft correlation into an explicit structural alignment. By integrating 2D Rotary Positional Embeddings with a Layout-Guided Attention Bias, StructMAR explicitly biases latent tokens toward their corresponding layout instances during attention computation. We further use Group Relative Policy Optimization (GRPO) as a final-stage policy refinement to reduce the mismatch between the MAR training objective and detector-based layout evaluation metrics. Evaluated on the COCO-Position and COCO-MIG benchmarks, StructMAR achieves state-of-the-art performance, reaching 57.2 AP and 79.4 mIoU on the former, and 61.7 ISR and 56.9 mIoU on the latter. These results, coupled with a 4.05$\times$ inference speedup, underscore the efficacy of explicit structural inductive biases in controllable autoregressive generation.
APA
Cao, G. & Zhang, J.. (2026). StructMAR: Structure-Aware Masked Autoregression for Explicit Layout Alignment in Text-to-Image Generation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:11720-11742 Available from https://proceedings.mlr.press/v306/cao26ab.html.

Related Material