MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

Chuanzhe Guo, Jingjing Wu, Sijun He, Yang Chen, Zhaoqi Kuang, Shilong Fan, Bingjin Chen, Siqi Bao, Jing Liu, Hua Wu, Qingfu Zhu, Wanxiang Che, Haifeng Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:38471-38495, 2026.

Abstract

The evolution of Large Language Model (LLM) agents for software engineering (SWE) is constrained by the scarcity of verifiable datasets, a bottleneck stemming from the complexity of constructing executable environments across diverse languages. To address this, we introduce MEnvAgent, a Multi-language framework for automated Environment construction that facilitates scalable generation of verifiable task instances. MEnvAgent employs a multi-agent Planning-Execution-Verification architecture to autonomously resolve construction failures and integrates a novel Environment Reuse Mechanism that reduces computational overhead by incrementally patching historical environments. Evaluations on MEnvBench, a new benchmark comprising 1,000 tasks across 10 languages, demonstrate that MEnvAgent outperforms baselines, improving Fail-to-Pass (F2P) rates by 8.6% while reducing time costs by 43%. Additionally, we demonstrate the utility of MEnvAgent by constructing MEnvData-SWE, the largest open-source polyglot dataset of realistic verifiable Docker environments to date, alongside solution trajectories that enable consistent performance gains on SWE tasks across a wide range of models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26aa, title = {{ME}nv{A}gent: Scalable Polyglot Environment Construction for Verifiable Software Engineering}, author = {Guo, Chuanzhe and Wu, Jingjing and He, Sijun and Chen, Yang and Kuang, Zhaoqi and Fan, Shilong and Chen, Bingjin and Bao, Siqi and Liu, Jing and Wu, Hua and Zhu, Qingfu and Che, Wanxiang and Wang, Haifeng}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {38471--38495}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26aa/guo26aa.pdf}, url = {https://proceedings.mlr.press/v306/guo26aa.html}, abstract = {The evolution of Large Language Model (LLM) agents for software engineering (SWE) is constrained by the scarcity of verifiable datasets, a bottleneck stemming from the complexity of constructing executable environments across diverse languages. To address this, we introduce MEnvAgent, a Multi-language framework for automated Environment construction that facilitates scalable generation of verifiable task instances. MEnvAgent employs a multi-agent Planning-Execution-Verification architecture to autonomously resolve construction failures and integrates a novel Environment Reuse Mechanism that reduces computational overhead by incrementally patching historical environments. Evaluations on MEnvBench, a new benchmark comprising 1,000 tasks across 10 languages, demonstrate that MEnvAgent outperforms baselines, improving Fail-to-Pass (F2P) rates by 8.6% while reducing time costs by 43%. Additionally, we demonstrate the utility of MEnvAgent by constructing MEnvData-SWE, the largest open-source polyglot dataset of realistic verifiable Docker environments to date, alongside solution trajectories that enable consistent performance gains on SWE tasks across a wide range of models.} }
Endnote
%0 Conference Paper %T MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering %A Chuanzhe Guo %A Jingjing Wu %A Sijun He %A Yang Chen %A Zhaoqi Kuang %A Shilong Fan %A Bingjin Chen %A Siqi Bao %A Jing Liu %A Hua Wu %A Qingfu Zhu %A Wanxiang Che %A Haifeng Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26aa %I PMLR %P 38471--38495 %U https://proceedings.mlr.press/v306/guo26aa.html %V 306 %X The evolution of Large Language Model (LLM) agents for software engineering (SWE) is constrained by the scarcity of verifiable datasets, a bottleneck stemming from the complexity of constructing executable environments across diverse languages. To address this, we introduce MEnvAgent, a Multi-language framework for automated Environment construction that facilitates scalable generation of verifiable task instances. MEnvAgent employs a multi-agent Planning-Execution-Verification architecture to autonomously resolve construction failures and integrates a novel Environment Reuse Mechanism that reduces computational overhead by incrementally patching historical environments. Evaluations on MEnvBench, a new benchmark comprising 1,000 tasks across 10 languages, demonstrate that MEnvAgent outperforms baselines, improving Fail-to-Pass (F2P) rates by 8.6% while reducing time costs by 43%. Additionally, we demonstrate the utility of MEnvAgent by constructing MEnvData-SWE, the largest open-source polyglot dataset of realistic verifiable Docker environments to date, alongside solution trajectories that enable consistent performance gains on SWE tasks across a wide range of models.
APA
Guo, C., Wu, J., He, S., Chen, Y., Kuang, Z., Fan, S., Chen, B., Bao, S., Liu, J., Wu, H., Zhu, Q., Che, W. & Wang, H.. (2026). MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:38471-38495 Available from https://proceedings.mlr.press/v306/guo26aa.html.

Related Material