From Diagrams to Code: Multilingual Programming with Visual Design

Linzheng Chai, Jian Yang, Shukai Liu, Wei Zhang, Liran Wang, Ke Jin, Tao Sun, Congnan Liu, Chenchen Zhang, Hualei Zhu, Jiaheng Liu, Xianjie Wu, Ge Zhang, Tianyu Liu, Zhoujun Li
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:12383-12423, 2026.

Abstract

In modern software development, particularly in emerging “vibe coding” paradigms, project implementation increasingly begins with visual interactions between users and AI coding assistants, where system architectures are communicated through visual designs before coding. This visual-first approach necessitates AI systems capable of interpreting diagrams across multiple programming languages. However, the development of such systems is severely hindered by the lack of large-scale multimodal training data and evaluation benchmarks. To address these limitations, we present M$^2$C-INSTRUCT, a comprehensive multilingual multimodal instruction-tuning dataset containing over 13.1M samples across 50+ programming languages, designed for visual understanding and diagram interpretation in code generation tasks. We validate our dataset by training M$^2$-CODER, a multilingual multimodal software developer that successfully integrates visual design inputs with textual instructions. We also introduce M$^2$EVAL, a novel multilingual evaluation benchmark for multimodal code generation performance. Experiments show our 7B M$^2$-CODER, performs on par with much larger 70B+ models, confirming the quality and effectiveness of our M$^2$C-INSTRUCT. Together, M$^2$C-INSTRUCT, M$^2$-CODER, and M$^2$EVAL provide essential infrastructure for visual-assisted programming in vibe-coding and visual-interactive development workflows.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chai26b, title = {From Diagrams to Code: Multilingual Programming with Visual Design}, author = {Chai, Linzheng and Yang, Jian and Liu, Shukai and Zhang, Wei and Wang, Liran and Jin, Ke and Sun, Tao and Liu, Congnan and Zhang, Chenchen and Zhu, Hualei and Liu, Jiaheng and Wu, Xianjie and Zhang, Ge and Liu, Tianyu and Li, Zhoujun}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {12383--12423}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chai26b/chai26b.pdf}, url = {https://proceedings.mlr.press/v306/chai26b.html}, abstract = {In modern software development, particularly in emerging “vibe coding” paradigms, project implementation increasingly begins with visual interactions between users and AI coding assistants, where system architectures are communicated through visual designs before coding. This visual-first approach necessitates AI systems capable of interpreting diagrams across multiple programming languages. However, the development of such systems is severely hindered by the lack of large-scale multimodal training data and evaluation benchmarks. To address these limitations, we present M$^2$C-INSTRUCT, a comprehensive multilingual multimodal instruction-tuning dataset containing over 13.1M samples across 50+ programming languages, designed for visual understanding and diagram interpretation in code generation tasks. We validate our dataset by training M$^2$-CODER, a multilingual multimodal software developer that successfully integrates visual design inputs with textual instructions. We also introduce M$^2$EVAL, a novel multilingual evaluation benchmark for multimodal code generation performance. Experiments show our 7B M$^2$-CODER, performs on par with much larger 70B+ models, confirming the quality and effectiveness of our M$^2$C-INSTRUCT. Together, M$^2$C-INSTRUCT, M$^2$-CODER, and M$^2$EVAL provide essential infrastructure for visual-assisted programming in vibe-coding and visual-interactive development workflows.} }
Endnote
%0 Conference Paper %T From Diagrams to Code: Multilingual Programming with Visual Design %A Linzheng Chai %A Jian Yang %A Shukai Liu %A Wei Zhang %A Liran Wang %A Ke Jin %A Tao Sun %A Congnan Liu %A Chenchen Zhang %A Hualei Zhu %A Jiaheng Liu %A Xianjie Wu %A Ge Zhang %A Tianyu Liu %A Zhoujun Li %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chai26b %I PMLR %P 12383--12423 %U https://proceedings.mlr.press/v306/chai26b.html %V 306 %X In modern software development, particularly in emerging “vibe coding” paradigms, project implementation increasingly begins with visual interactions between users and AI coding assistants, where system architectures are communicated through visual designs before coding. This visual-first approach necessitates AI systems capable of interpreting diagrams across multiple programming languages. However, the development of such systems is severely hindered by the lack of large-scale multimodal training data and evaluation benchmarks. To address these limitations, we present M$^2$C-INSTRUCT, a comprehensive multilingual multimodal instruction-tuning dataset containing over 13.1M samples across 50+ programming languages, designed for visual understanding and diagram interpretation in code generation tasks. We validate our dataset by training M$^2$-CODER, a multilingual multimodal software developer that successfully integrates visual design inputs with textual instructions. We also introduce M$^2$EVAL, a novel multilingual evaluation benchmark for multimodal code generation performance. Experiments show our 7B M$^2$-CODER, performs on par with much larger 70B+ models, confirming the quality and effectiveness of our M$^2$C-INSTRUCT. Together, M$^2$C-INSTRUCT, M$^2$-CODER, and M$^2$EVAL provide essential infrastructure for visual-assisted programming in vibe-coding and visual-interactive development workflows.
APA
Chai, L., Yang, J., Liu, S., Zhang, W., Wang, L., Jin, K., Sun, T., Liu, C., Zhang, C., Zhu, H., Liu, J., Wu, X., Zhang, G., Liu, T. & Li, Z.. (2026). From Diagrams to Code: Multilingual Programming with Visual Design. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:12383-12423 Available from https://proceedings.mlr.press/v306/chai26b.html.

Related Material