OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yifan Zhang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:13490-13512, 2026.

Abstract

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for physics/chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18% on editing tasks UniWorld-V1 on ImgEdit-Bench and 13% on generation tasks Harmon on GenEval. Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26c, title = {{O}pen{GPT}-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing}, author = {Chen, Zhihong and Bai, Xuehai and Shi, Yang and Fu, Chaoyou and Zhang, Huanyu and Wang, Haotian and Sun, Xiaoyan and Zhang, Zhang and Wang, Liang and Zhang, Yuanxing and Wan, Pengfei and Zhang, Yifan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {13490--13512}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26c/chen26c.pdf}, url = {https://proceedings.mlr.press/v306/chen26c.html}, abstract = {The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for physics/chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18% on editing tasks UniWorld-V1 on ImgEdit-Bench and 13% on generation tasks Harmon on GenEval. Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.} }
Endnote
%0 Conference Paper %T OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing %A Zhihong Chen %A Xuehai Bai %A Yang Shi %A Chaoyou Fu %A Huanyu Zhang %A Haotian Wang %A Xiaoyan Sun %A Zhang Zhang %A Liang Wang %A Yuanxing Zhang %A Pengfei Wan %A Yifan Zhang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26c %I PMLR %P 13490--13512 %U https://proceedings.mlr.press/v306/chen26c.html %V 306 %X The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for physics/chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18% on editing tasks UniWorld-V1 on ImgEdit-Bench and 13% on generation tasks Harmon on GenEval. Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.
APA
Chen, Z., Bai, X., Shi, Y., Fu, C., Zhang, H., Wang, H., Sun, X., Zhang, Z., Wang, L., Zhang, Y., Wan, P. & Zhang, Y.. (2026). OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:13490-13512 Available from https://proceedings.mlr.press/v306/chen26c.html.

Related Material