AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Junzhi Chen, Harsh Trivedi, Jane Pan, Michael Jq Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16996-17023, 2026.

Abstract

Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a “user-in-the-loop” benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark’s difficulty and its potential to advance research on user-in-the-loop tool-use agents.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26eg, title = {{A}pp{W}orld-{UL}: Benchmarking Diverse Agent-User Interactions for Tool-Use}, author = {Chen, Junzhi and Trivedi, Harsh and Pan, Jane and Zhang, Michael Jq and Srinivasan, Tejas and Balasubramanian, Niranjan and Sabharwal, Ashish}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16996--17023}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26eg/chen26eg.pdf}, url = {https://proceedings.mlr.press/v306/chen26eg.html}, abstract = {Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a “user-in-the-loop” benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark’s difficulty and its potential to advance research on user-in-the-loop tool-use agents.} }
Endnote
%0 Conference Paper %T AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use %A Junzhi Chen %A Harsh Trivedi %A Jane Pan %A Michael Jq Zhang %A Tejas Srinivasan %A Niranjan Balasubramanian %A Ashish Sabharwal %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26eg %I PMLR %P 16996--17023 %U https://proceedings.mlr.press/v306/chen26eg.html %V 306 %X Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a “user-in-the-loop” benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark’s difficulty and its potential to advance research on user-in-the-loop tool-use agents.
APA
Chen, J., Trivedi, H., Pan, J., Zhang, M.J., Srinivasan, T., Balasubramanian, N. & Sabharwal, A.. (2026). AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16996-17023 Available from https://proceedings.mlr.press/v306/chen26eg.html.

Related Material