CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments

Lingyue Fu, Xin Ding, Linyue Pan, Yaoming Zhu, Shao Zhang, Lin Qiu, Xuezhi Cao, Xunliang Cai, Jiaxin Ding, Weiwen Liu, Weinan Zhang, Yong Yu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:31726-31755, 2026.

Abstract

Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent’s capability for continuous code optimization and multi-turn iterative development. To bridge this gap, we introduce CATArena, a framework designed to evaluate the evolutionary capabilities of code agents via iterative tournaments. Agents engage in multi-turn tournaments and continuously refine their code through self-reflection and peer-learning based on comprehensive execution feedback. For evaluation, we propose a dual-metric system to decouple static generation proficiency from evolutionary potential. Extensive experiments reveal that an agent’s evolutionary potential is not strictly correlated with its initial proficiency. Our analysis further reveals that current agents struggle to concurrently leverage both peer-learning and self-reflection for effective performance gains. Furthermore, the results validate CATArena’s high extensibility and resistance to variance tasks, establishing it as a continuous and reliable standard for assessing the evolutionary capability of LLM code agents.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fu26a, title = {{CATA}rena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments}, author = {Fu, Lingyue and Ding, Xin and Pan, Linyue and Zhu, Yaoming and Zhang, Shao and Qiu, Lin and Cao, Xuezhi and Cai, Xunliang and Ding, Jiaxin and Liu, Weiwen and Zhang, Weinan and Yu, Yong}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {31726--31755}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fu26a/fu26a.pdf}, url = {https://proceedings.mlr.press/v306/fu26a.html}, abstract = {Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent’s capability for continuous code optimization and multi-turn iterative development. To bridge this gap, we introduce CATArena, a framework designed to evaluate the evolutionary capabilities of code agents via iterative tournaments. Agents engage in multi-turn tournaments and continuously refine their code through self-reflection and peer-learning based on comprehensive execution feedback. For evaluation, we propose a dual-metric system to decouple static generation proficiency from evolutionary potential. Extensive experiments reveal that an agent’s evolutionary potential is not strictly correlated with its initial proficiency. Our analysis further reveals that current agents struggle to concurrently leverage both peer-learning and self-reflection for effective performance gains. Furthermore, the results validate CATArena’s high extensibility and resistance to variance tasks, establishing it as a continuous and reliable standard for assessing the evolutionary capability of LLM code agents.} }
Endnote
%0 Conference Paper %T CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments %A Lingyue Fu %A Xin Ding %A Linyue Pan %A Yaoming Zhu %A Shao Zhang %A Lin Qiu %A Xuezhi Cao %A Xunliang Cai %A Jiaxin Ding %A Weiwen Liu %A Weinan Zhang %A Yong Yu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fu26a %I PMLR %P 31726--31755 %U https://proceedings.mlr.press/v306/fu26a.html %V 306 %X Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent’s capability for continuous code optimization and multi-turn iterative development. To bridge this gap, we introduce CATArena, a framework designed to evaluate the evolutionary capabilities of code agents via iterative tournaments. Agents engage in multi-turn tournaments and continuously refine their code through self-reflection and peer-learning based on comprehensive execution feedback. For evaluation, we propose a dual-metric system to decouple static generation proficiency from evolutionary potential. Extensive experiments reveal that an agent’s evolutionary potential is not strictly correlated with its initial proficiency. Our analysis further reveals that current agents struggle to concurrently leverage both peer-learning and self-reflection for effective performance gains. Furthermore, the results validate CATArena’s high extensibility and resistance to variance tasks, establishing it as a continuous and reliable standard for assessing the evolutionary capability of LLM code agents.
APA
Fu, L., Ding, X., Pan, L., Zhu, Y., Zhang, S., Qiu, L., Cao, X., Cai, X., Ding, J., Liu, W., Zhang, W. & Yu, Y.. (2026). CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:31726-31755 Available from https://proceedings.mlr.press/v306/fu26a.html.

Related Material