How can we assess human-agent interactions? Case studies in software agent design

Valerie Chen, Rohit Malhotra, Xingyao Wang, Juan Michelini, Xuhui Zhou, Aditya Bharat Soni, Hoang H. Tran, Calvin Smith, Ameet Talwalkar, Graham Neubig
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16105-16124, 2026.

Abstract

While benchmarks measure the accuracy of LLM-powered agents, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy PULSE n software engineering—one of the highest-impact, real-world domains for LLM agents—via a large-scale web platform built around the open-source agent OpenHands. Across 15k users, we evaluate how three agent design decisions impact developer satisfaction rates. We also show how PULSE can lead to more robust conclusions about agent design, reducing confidence intervals by 40% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results with benchmark performance (e.g., the anti-correlation between claude-sonnet-4 and gpt-5, underscoring the limitations of benchmark-driven evaluation. Our framework PULSE provides guidance for future evaluations, and our findings identify opportunities for better software agent designs.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26da, title = {How can we assess human-agent interactions? {C}ase studies in software agent design}, author = {Chen, Valerie and Malhotra, Rohit and Wang, Xingyao and Michelini, Juan and Zhou, Xuhui and Soni, Aditya Bharat and Tran, Hoang H. and Smith, Calvin and Talwalkar, Ameet and Neubig, Graham}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16105--16124}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26da/chen26da.pdf}, url = {https://proceedings.mlr.press/v306/chen26da.html}, abstract = {While benchmarks measure the accuracy of LLM-powered agents, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy PULSE n software engineering—one of the highest-impact, real-world domains for LLM agents—via a large-scale web platform built around the open-source agent OpenHands. Across 15k users, we evaluate how three agent design decisions impact developer satisfaction rates. We also show how PULSE can lead to more robust conclusions about agent design, reducing confidence intervals by 40% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results with benchmark performance (e.g., the anti-correlation between claude-sonnet-4 and gpt-5, underscoring the limitations of benchmark-driven evaluation. Our framework PULSE provides guidance for future evaluations, and our findings identify opportunities for better software agent designs.} }
Endnote
%0 Conference Paper %T How can we assess human-agent interactions? Case studies in software agent design %A Valerie Chen %A Rohit Malhotra %A Xingyao Wang %A Juan Michelini %A Xuhui Zhou %A Aditya Bharat Soni %A Hoang H. Tran %A Calvin Smith %A Ameet Talwalkar %A Graham Neubig %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26da %I PMLR %P 16105--16124 %U https://proceedings.mlr.press/v306/chen26da.html %V 306 %X While benchmarks measure the accuracy of LLM-powered agents, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases. In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. First, we propose PULSE, a framework for more efficient human-centric evaluation of agent designs, which comprises collecting user feedback, training an ML model to predict user satisfaction, and computing results by combining human satisfaction ratings with model-generated pseudo-labels. Second, we deploy PULSE n software engineering—one of the highest-impact, real-world domains for LLM agents—via a large-scale web platform built around the open-source agent OpenHands. Across 15k users, we evaluate how three agent design decisions impact developer satisfaction rates. We also show how PULSE can lead to more robust conclusions about agent design, reducing confidence intervals by 40% compared to a standard A/B test. Finally, we find substantial discrepancies between in-the-wild results with benchmark performance (e.g., the anti-correlation between claude-sonnet-4 and gpt-5, underscoring the limitations of benchmark-driven evaluation. Our framework PULSE provides guidance for future evaluations, and our findings identify opportunities for better software agent designs.
APA
Chen, V., Malhotra, R., Wang, X., Michelini, J., Zhou, X., Soni, A.B., Tran, H.H., Smith, C., Talwalkar, A. & Neubig, G.. (2026). How can we assess human-agent interactions? Case studies in software agent design. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16105-16124 Available from https://proceedings.mlr.press/v306/chen26da.html.

Related Material