InteractComp: Evaluating Search Agents With Ambiguous Queries

Mingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong, Jiayi Zhang, Fashen Ren, Jinyi Bai, Fuzhen Yang, Dayi Miao, Zhaoyang Yu, Yifan Wu, Yanfei Zhang, Fengwei Teng, Yingjia Wan, Song Hu, Yude Li, Xin Jin, Conghao Hu, Haoyu Li, Qirui Fu, Tai Zhong, Xinyu Wang, Xiangru Tang, Nan Tang, Chenglin Wu, Yuyu Luo
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:23953-23972, 2026.

Abstract

Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a practical failure mode: agents may face ambiguous requests where the intended target cannot be identified without clarification. Yet most agents lack interactive mechanisms during the search process, and existing benchmarks cannot assess this capability. To address this gap, we introduce InteractComp, a benchmark designed to evaluate whether search agents can recognize query ambiguity and actively interact to resolve it during search. Following the principle of easy to verify, interact to disambiguate, we construct 210 expert-curated questions across 9 domains through a target-distractor methodology that creates controlled ambiguity resolvable only through interaction. Evaluation of 17 models reveals striking failure: the best model achieves only 13.73% accuracy despite 71.50% with complete context, exposing systematic overconfidence rather than reasoning deficits. Forced interaction produces dramatic gains, demonstrating latent capability current strategies fail to engage. Longitudinal analysis shows interaction capabilities stagnated over 15 months while search performance improved seven-fold, revealing a critical blind spot. This stagnation, coupled with the immediate feedback inherent to search tasks, makes InteractComp a valuable resource for both evaluating and training interaction capabilities in search agents. The code is available at https://github.com/FoundationAgents/InteractComp

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-deng26h, title = {{I}nteract{C}omp: Evaluating Search Agents With Ambiguous Queries}, author = {Deng, Mingyi and Huang, Lijun and Fan, Yani and Kong, Fanqi and Zhang, Jiayi and Ren, Fashen and Bai, Jinyi and Yang, Fuzhen and Miao, Dayi and Yu, Zhaoyang and Wu, Yifan and Zhang, Yanfei and Teng, Fengwei and Wan, Yingjia and Hu, Song and Li, Yude and Jin, Xin and Hu, Conghao and Li, Haoyu and Fu, Qirui and Zhong, Tai and Wang, Xinyu and Tang, Xiangru and Tang, Nan and Wu, Chenglin and Luo, Yuyu}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {23953--23972}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/deng26h/deng26h.pdf}, url = {https://proceedings.mlr.press/v306/deng26h.html}, abstract = {Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a practical failure mode: agents may face ambiguous requests where the intended target cannot be identified without clarification. Yet most agents lack interactive mechanisms during the search process, and existing benchmarks cannot assess this capability. To address this gap, we introduce InteractComp, a benchmark designed to evaluate whether search agents can recognize query ambiguity and actively interact to resolve it during search. Following the principle of easy to verify, interact to disambiguate, we construct 210 expert-curated questions across 9 domains through a target-distractor methodology that creates controlled ambiguity resolvable only through interaction. Evaluation of 17 models reveals striking failure: the best model achieves only 13.73% accuracy despite 71.50% with complete context, exposing systematic overconfidence rather than reasoning deficits. Forced interaction produces dramatic gains, demonstrating latent capability current strategies fail to engage. Longitudinal analysis shows interaction capabilities stagnated over 15 months while search performance improved seven-fold, revealing a critical blind spot. This stagnation, coupled with the immediate feedback inherent to search tasks, makes InteractComp a valuable resource for both evaluating and training interaction capabilities in search agents. The code is available at https://github.com/FoundationAgents/InteractComp} }
Endnote
%0 Conference Paper %T InteractComp: Evaluating Search Agents With Ambiguous Queries %A Mingyi Deng %A Lijun Huang %A Yani Fan %A Fanqi Kong %A Jiayi Zhang %A Fashen Ren %A Jinyi Bai %A Fuzhen Yang %A Dayi Miao %A Zhaoyang Yu %A Yifan Wu %A Yanfei Zhang %A Fengwei Teng %A Yingjia Wan %A Song Hu %A Yude Li %A Xin Jin %A Conghao Hu %A Haoyu Li %A Qirui Fu %A Tai Zhong %A Xinyu Wang %A Xiangru Tang %A Nan Tang %A Chenglin Wu %A Yuyu Luo %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-deng26h %I PMLR %P 23953--23972 %U https://proceedings.mlr.press/v306/deng26h.html %V 306 %X Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a practical failure mode: agents may face ambiguous requests where the intended target cannot be identified without clarification. Yet most agents lack interactive mechanisms during the search process, and existing benchmarks cannot assess this capability. To address this gap, we introduce InteractComp, a benchmark designed to evaluate whether search agents can recognize query ambiguity and actively interact to resolve it during search. Following the principle of easy to verify, interact to disambiguate, we construct 210 expert-curated questions across 9 domains through a target-distractor methodology that creates controlled ambiguity resolvable only through interaction. Evaluation of 17 models reveals striking failure: the best model achieves only 13.73% accuracy despite 71.50% with complete context, exposing systematic overconfidence rather than reasoning deficits. Forced interaction produces dramatic gains, demonstrating latent capability current strategies fail to engage. Longitudinal analysis shows interaction capabilities stagnated over 15 months while search performance improved seven-fold, revealing a critical blind spot. This stagnation, coupled with the immediate feedback inherent to search tasks, makes InteractComp a valuable resource for both evaluating and training interaction capabilities in search agents. The code is available at https://github.com/FoundationAgents/InteractComp
APA
Deng, M., Huang, L., Fan, Y., Kong, F., Zhang, J., Ren, F., Bai, J., Yang, F., Miao, D., Yu, Z., Wu, Y., Zhang, Y., Teng, F., Wan, Y., Hu, S., Li, Y., Jin, X., Hu, C., Li, H., Fu, Q., Zhong, T., Wang, X., Tang, X., Tang, N., Wu, C. & Luo, Y.. (2026). InteractComp: Evaluating Search Agents With Ambiguous Queries. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:23953-23972 Available from https://proceedings.mlr.press/v306/deng26h.html.

Related Material