Measuring Agents in Production

Melissa Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi, Xiaoyuan Liu, Jared Quincy Davis, Emmanuele Lacavalla, Alessandro Basile, Shuyi Yang, Paul Castro, Daniel Kang, Koushik Sen, Dawn Song, Joseph E. Gonzalez, Ion Stoica, Matei Zaharia, Marquita Ellis
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:95649-95695, 2026.

Abstract

LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first systematic study of Measuring Agents in Production, MAP, using first-hand data from agent developers. We conducted 20 case studies via in-depth interviews and surveyed 86 deployed systems practitioners across 26 domains. We investigate why organizations build agents, how they build them, how they evaluate them, and their top development challenges. Our study finds that production agents are built using simple, controllable approaches: 68% execute at most 10 steps before human intervention, 70% rely on prompting off-the-shelf models instead of weight tuning, and 74% depend primarily on human evaluation. Reliability (consistent correct behavior over time) remains the top development challenge, which practitioners currently address through systems-level design. MAP documents the current state of production agents, providing the research community with visibility into deployment realities and underexplored research avenues.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-pan26a, title = {Measuring Agents in Production}, author = {Pan, Melissa and Arabzadeh, Negar and Cogo, Riccardo and Zhu, Yuxuan and Xiong, Alexander and A Agrawal, Lakshya and Mao, Huanzhi and Shen, Emma and Pallerla, Sid and Patel, Liana and Liu, Shu and Shi, Tianneng and Liu, Xiaoyuan and Davis, Jared Quincy and Lacavalla, Emmanuele and Basile, Alessandro and Yang, Shuyi and Castro, Paul and Kang, Daniel and Sen, Koushik and Song, Dawn and Gonzalez, Joseph E. and Stoica, Ion and Zaharia, Matei and Ellis, Marquita}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {95649--95695}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/pan26a/pan26a.pdf}, url = {https://proceedings.mlr.press/v306/pan26a.html}, abstract = {LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first systematic study of Measuring Agents in Production, MAP, using first-hand data from agent developers. We conducted 20 case studies via in-depth interviews and surveyed 86 deployed systems practitioners across 26 domains. We investigate why organizations build agents, how they build them, how they evaluate them, and their top development challenges. Our study finds that production agents are built using simple, controllable approaches: 68% execute at most 10 steps before human intervention, 70% rely on prompting off-the-shelf models instead of weight tuning, and 74% depend primarily on human evaluation. Reliability (consistent correct behavior over time) remains the top development challenge, which practitioners currently address through systems-level design. MAP documents the current state of production agents, providing the research community with visibility into deployment realities and underexplored research avenues.} }
Endnote
%0 Conference Paper %T Measuring Agents in Production %A Melissa Pan %A Negar Arabzadeh %A Riccardo Cogo %A Yuxuan Zhu %A Alexander Xiong %A Lakshya A Agrawal %A Huanzhi Mao %A Emma Shen %A Sid Pallerla %A Liana Patel %A Shu Liu %A Tianneng Shi %A Xiaoyuan Liu %A Jared Quincy Davis %A Emmanuele Lacavalla %A Alessandro Basile %A Shuyi Yang %A Paul Castro %A Daniel Kang %A Koushik Sen %A Dawn Song %A Joseph E. Gonzalez %A Ion Stoica %A Matei Zaharia %A Marquita Ellis %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-pan26a %I PMLR %P 95649--95695 %U https://proceedings.mlr.press/v306/pan26a.html %V 306 %X LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first systematic study of Measuring Agents in Production, MAP, using first-hand data from agent developers. We conducted 20 case studies via in-depth interviews and surveyed 86 deployed systems practitioners across 26 domains. We investigate why organizations build agents, how they build them, how they evaluate them, and their top development challenges. Our study finds that production agents are built using simple, controllable approaches: 68% execute at most 10 steps before human intervention, 70% rely on prompting off-the-shelf models instead of weight tuning, and 74% depend primarily on human evaluation. Reliability (consistent correct behavior over time) remains the top development challenge, which practitioners currently address through systems-level design. MAP documents the current state of production agents, providing the research community with visibility into deployment realities and underexplored research avenues.
APA
Pan, M., Arabzadeh, N., Cogo, R., Zhu, Y., Xiong, A., A Agrawal, L., Mao, H., Shen, E., Pallerla, S., Patel, L., Liu, S., Shi, T., Liu, X., Davis, J.Q., Lacavalla, E., Basile, A., Yang, S., Castro, P., Kang, D., Sen, K., Song, D., Gonzalez, J.E., Stoica, I., Zaharia, M. & Ellis, M.. (2026). Measuring Agents in Production. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:95649-95695 Available from https://proceedings.mlr.press/v306/pan26a.html.

Related Material