GameDevBench: Evaluating Agentic Capabilities Through Game Development

Wayne Chi, Yixiong Fang, Arnav Yayavaram, Siddharth Yayavaram, Seth Karten, Qiuhong Anna Wei, Runkun Chen, Alexander Wang, Valerie Chen, Ameet Talwalkar, Chris Donahue
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:19395-19423, 2026.

Abstract

Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for deep multimodal understanding. In game development, agents must navigate large, dense codebases while manipulating intrinsically multimodal assets such as shaders, sprites, and animations within a visual game scene. We present GameDevBench, the first benchmark for evaluating agents on game development tasks. GameDevBench consists of 333 tasks derived from web and video tutorials. Tasks require significant multimodal understanding and are complex—the average solution requires over three times the lines of code and file changes compared to prior software development benchmarks. Agents struggle with game development, with the best agent and method solving only $53.8%$ of tasks. We find a strong correlation between perceived task difficulty and multimodal complexity, with average success rate dropping from $51.4%$ on gameplay-oriented tasks to $33.0%$ on 2D graphics tasks. To improve multimodal capability, we introduce two simple image and video-based feedback mechanisms for agents. Despite their simplicity, these methods consistently improve performance, increasing GPT-5.4’s performance from $41.1%$ to $52.0%$ when given visual feedback.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chi26a, title = {{G}ame{D}ev{B}ench: Evaluating Agentic Capabilities Through Game Development}, author = {Chi, Wayne and Fang, Yixiong and Yayavaram, Arnav and Yayavaram, Siddharth and Karten, Seth and Wei, Qiuhong Anna and Chen, Runkun and Wang, Alexander and Chen, Valerie and Talwalkar, Ameet and Donahue, Chris}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {19395--19423}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chi26a/chi26a.pdf}, url = {https://proceedings.mlr.press/v306/chi26a.html}, abstract = {Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for deep multimodal understanding. In game development, agents must navigate large, dense codebases while manipulating intrinsically multimodal assets such as shaders, sprites, and animations within a visual game scene. We present GameDevBench, the first benchmark for evaluating agents on game development tasks. GameDevBench consists of 333 tasks derived from web and video tutorials. Tasks require significant multimodal understanding and are complex—the average solution requires over three times the lines of code and file changes compared to prior software development benchmarks. Agents struggle with game development, with the best agent and method solving only $53.8%$ of tasks. We find a strong correlation between perceived task difficulty and multimodal complexity, with average success rate dropping from $51.4%$ on gameplay-oriented tasks to $33.0%$ on 2D graphics tasks. To improve multimodal capability, we introduce two simple image and video-based feedback mechanisms for agents. Despite their simplicity, these methods consistently improve performance, increasing GPT-5.4’s performance from $41.1%$ to $52.0%$ when given visual feedback.} }
Endnote
%0 Conference Paper %T GameDevBench: Evaluating Agentic Capabilities Through Game Development %A Wayne Chi %A Yixiong Fang %A Arnav Yayavaram %A Siddharth Yayavaram %A Seth Karten %A Qiuhong Anna Wei %A Runkun Chen %A Alexander Wang %A Valerie Chen %A Ameet Talwalkar %A Chris Donahue %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chi26a %I PMLR %P 19395--19423 %U https://proceedings.mlr.press/v306/chi26a.html %V 306 %X Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for deep multimodal understanding. In game development, agents must navigate large, dense codebases while manipulating intrinsically multimodal assets such as shaders, sprites, and animations within a visual game scene. We present GameDevBench, the first benchmark for evaluating agents on game development tasks. GameDevBench consists of 333 tasks derived from web and video tutorials. Tasks require significant multimodal understanding and are complex—the average solution requires over three times the lines of code and file changes compared to prior software development benchmarks. Agents struggle with game development, with the best agent and method solving only $53.8%$ of tasks. We find a strong correlation between perceived task difficulty and multimodal complexity, with average success rate dropping from $51.4%$ on gameplay-oriented tasks to $33.0%$ on 2D graphics tasks. To improve multimodal capability, we introduce two simple image and video-based feedback mechanisms for agents. Despite their simplicity, these methods consistently improve performance, increasing GPT-5.4’s performance from $41.1%$ to $52.0%$ when given visual feedback.
APA
Chi, W., Fang, Y., Yayavaram, A., Yayavaram, S., Karten, S., Wei, Q.A., Chen, R., Wang, A., Chen, V., Talwalkar, A. & Donahue, C.. (2026). GameDevBench: Evaluating Agentic Capabilities Through Game Development. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:19395-19423 Available from https://proceedings.mlr.press/v306/chi26a.html.

Related Material