NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents

Jingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang, Zhou Huan, Hongwan Gao, Xiang Gao, Chao He, Yue Hou, Fei Hu, Zhaojian Li, Weiran Shi, Zaiyuan Wang, Daoguang Zan, Chenchen Zhang, Xiaoxu Zhang, Chen Qizhi, Xianfu Cheng, Bo Deng, Qingshui Gu, Kai Hua, Juntao Lin, Pai Liu, Mingchen Li, Minghao Li, Xuanguang Pan, Zifan Peng, Yujia Qin, Yong Shan, Zhewen Tan, Haoran Wang, Zihan Wang, Weihao Xie, Yishuo Yuan, Jiayu Zhang, Yunfei Zhao, He Zhu, Liya Zhu, Chenyang Zou, Ming Ding, Jianpeng Jiao, Jiaheng Liu, Minghao Liu, Qian Liu, Chongyang Tao, Jian Yang, Tong Yang, Zhaoxiang Zhang, Xinjie Chen, Wenhao Huang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:25035-25054, 2026.

Abstract

Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. Experiments across state-of-the-art open- and closed-source models reveal that long-horizon repository generation remains largely unsolved, with even the strongest agents achieving merely 40% average test pass rates and rarely completing an entire repository correctly. Further analysis identifies systematic long-horizon failure modes, including premature termination, loss of global coherence, fragile cross-file dependencies, and inadequate planning over hundreds of interaction steps. These results position NL2Repo-Bench as a rigorous, execution-based testbed for evaluating sustained agentic competence and highlight long-horizon reasoning as a key bottleneck for autonomous coding agents. Our data and code are available at https://github.com/multimodal-art-projection/NL2RepoBench.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ding26j, title = {{NL}2{R}epo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents}, author = {Ding, Jingzhe and Long, Shengda and Pu, Changxin and Zhang, Ge and Huan, Zhou and Gao, Hongwan and Gao, Xiang and He, Chao and Hou, Yue and Hu, Fei and Li, Zhaojian and Shi, Weiran and Wang, Zaiyuan and Zan, Daoguang and Zhang, Chenchen and Zhang, Xiaoxu and Qizhi, Chen and Cheng, Xianfu and Deng, Bo and Gu, Qingshui and Hua, Kai and Lin, Juntao and Liu, Pai and Li, Mingchen and Li, Minghao and Pan, Xuanguang and Peng, Zifan and Qin, Yujia and Shan, Yong and Tan, Zhewen and Wang, Haoran and Wang, Zihan and Xie, Weihao and Yuan, Yishuo and Zhang, Jiayu and Zhao, Yunfei and Zhu, He and Zhu, Liya and Zou, Chenyang and Ding, Ming and Jiao, Jianpeng and Liu, Jiaheng and Liu, Minghao and Liu, Qian and Tao, Chongyang and Yang, Jian and Yang, Tong and Zhang, Zhaoxiang and Chen, Xinjie and Huang, Wenhao}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {25035--25054}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ding26j/ding26j.pdf}, url = {https://proceedings.mlr.press/v306/ding26j.html}, abstract = {Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. Experiments across state-of-the-art open- and closed-source models reveal that long-horizon repository generation remains largely unsolved, with even the strongest agents achieving merely 40% average test pass rates and rarely completing an entire repository correctly. Further analysis identifies systematic long-horizon failure modes, including premature termination, loss of global coherence, fragile cross-file dependencies, and inadequate planning over hundreds of interaction steps. These results position NL2Repo-Bench as a rigorous, execution-based testbed for evaluating sustained agentic competence and highlight long-horizon reasoning as a key bottleneck for autonomous coding agents. Our data and code are available at https://github.com/multimodal-art-projection/NL2RepoBench.} }
Endnote
%0 Conference Paper %T NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents %A Jingzhe Ding %A Shengda Long %A Changxin Pu %A Ge Zhang %A Zhou Huan %A Hongwan Gao %A Xiang Gao %A Chao He %A Yue Hou %A Fei Hu %A Zhaojian Li %A Weiran Shi %A Zaiyuan Wang %A Daoguang Zan %A Chenchen Zhang %A Xiaoxu Zhang %A Chen Qizhi %A Xianfu Cheng %A Bo Deng %A Qingshui Gu %A Kai Hua %A Juntao Lin %A Pai Liu %A Mingchen Li %A Minghao Li %A Xuanguang Pan %A Zifan Peng %A Yujia Qin %A Yong Shan %A Zhewen Tan %A Haoran Wang %A Zihan Wang %A Weihao Xie %A Yishuo Yuan %A Jiayu Zhang %A Yunfei Zhao %A He Zhu %A Liya Zhu %A Chenyang Zou %A Ming Ding %A Jianpeng Jiao %A Jiaheng Liu %A Minghao Liu %A Qian Liu %A Chongyang Tao %A Jian Yang %A Tong Yang %A Zhaoxiang Zhang %A Xinjie Chen %A Wenhao Huang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ding26j %I PMLR %P 25035--25054 %U https://proceedings.mlr.press/v306/ding26j.html %V 306 %X Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. Experiments across state-of-the-art open- and closed-source models reveal that long-horizon repository generation remains largely unsolved, with even the strongest agents achieving merely 40% average test pass rates and rarely completing an entire repository correctly. Further analysis identifies systematic long-horizon failure modes, including premature termination, loss of global coherence, fragile cross-file dependencies, and inadequate planning over hundreds of interaction steps. These results position NL2Repo-Bench as a rigorous, execution-based testbed for evaluating sustained agentic competence and highlight long-horizon reasoning as a key bottleneck for autonomous coding agents. Our data and code are available at https://github.com/multimodal-art-projection/NL2RepoBench.
APA
Ding, J., Long, S., Pu, C., Zhang, G., Huan, Z., Gao, H., Gao, X., He, C., Hou, Y., Hu, F., Li, Z., Shi, W., Wang, Z., Zan, D., Zhang, C., Zhang, X., Qizhi, C., Cheng, X., Deng, B., Gu, Q., Hua, K., Lin, J., Liu, P., Li, M., Li, M., Pan, X., Peng, Z., Qin, Y., Shan, Y., Tan, Z., Wang, H., Wang, Z., Xie, W., Yuan, Y., Zhang, J., Zhao, Y., Zhu, H., Zhu, L., Zou, C., Ding, M., Jiao, J., Liu, J., Liu, M., Liu, Q., Tao, C., Yang, J., Yang, T., Zhang, Z., Chen, X. & Huang, W.. (2026). NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:25035-25054 Available from https://proceedings.mlr.press/v306/ding26j.html.

Related Material