Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

Łukasz Borchmann, Jordy Van Landeghem, Michał Turski, Shreyansh Padarha, Ryan Othniel Kearns, Adam Mahdi, Niels Rogge, Clémentine Fourrier, Siwei Han, Huaxiu Yao, Artemis Llabrés, Yiming Xu, Dimosthenis Karatzas, Hao Zhang, Anupam Datta
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:9178-9241, 2026.

Abstract

Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoning, or merely stochastic trial-and-error search? To address this, we introduce MADQA, a benchmark of 2,250 human-authored questions grounded in 800 heterogeneous PDF documents. Guided by Classical Test Theory, we design it to maximize discriminative power across varying levels of agentic abilities. To evaluate agentic behavior, we introduce a novel protocol that measures the accuracy-effort trade-off. Using this framework, we show that while the best agents can match human searchers in raw accuracy, they succeed on largely different questions and rely on brute-force search to compensate for weak strategic planning. They fail to close the nearly 20% gap to oracle performance, persisting in unproductive loops. We release the dataset and evaluation harness to help facilitate the transition from brute-force retrieval to calibrated, efficient reasoning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-borchmann26a, title = {Strategic Navigation or Stochastic Search? {H}ow Agents and Humans Reason Over Document Collections}, author = {Borchmann, {\L}ukasz and Van Landeghem, Jordy and Turski, Micha{\l} and Padarha, Shreyansh and Kearns, Ryan Othniel and Mahdi, Adam and Rogge, Niels and Fourrier, Cl\'{e}mentine and Han, Siwei and Yao, Huaxiu and Llabr\'{e}s, Artemis and Xu, Yiming and Karatzas, Dimosthenis and Zhang, Hao and Datta, Anupam}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {9178--9241}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/borchmann26a/borchmann26a.pdf}, url = {https://proceedings.mlr.press/v306/borchmann26a.html}, abstract = {Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoning, or merely stochastic trial-and-error search? To address this, we introduce MADQA, a benchmark of 2,250 human-authored questions grounded in 800 heterogeneous PDF documents. Guided by Classical Test Theory, we design it to maximize discriminative power across varying levels of agentic abilities. To evaluate agentic behavior, we introduce a novel protocol that measures the accuracy-effort trade-off. Using this framework, we show that while the best agents can match human searchers in raw accuracy, they succeed on largely different questions and rely on brute-force search to compensate for weak strategic planning. They fail to close the nearly 20% gap to oracle performance, persisting in unproductive loops. We release the dataset and evaluation harness to help facilitate the transition from brute-force retrieval to calibrated, efficient reasoning.} }
Endnote
%0 Conference Paper %T Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections %A Łukasz Borchmann %A Jordy Van Landeghem %A Michał Turski %A Shreyansh Padarha %A Ryan Othniel Kearns %A Adam Mahdi %A Niels Rogge %A Clémentine Fourrier %A Siwei Han %A Huaxiu Yao %A Artemis Llabrés %A Yiming Xu %A Dimosthenis Karatzas %A Hao Zhang %A Anupam Datta %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-borchmann26a %I PMLR %P 9178--9241 %U https://proceedings.mlr.press/v306/borchmann26a.html %V 306 %X Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoning, or merely stochastic trial-and-error search? To address this, we introduce MADQA, a benchmark of 2,250 human-authored questions grounded in 800 heterogeneous PDF documents. Guided by Classical Test Theory, we design it to maximize discriminative power across varying levels of agentic abilities. To evaluate agentic behavior, we introduce a novel protocol that measures the accuracy-effort trade-off. Using this framework, we show that while the best agents can match human searchers in raw accuracy, they succeed on largely different questions and rely on brute-force search to compensate for weak strategic planning. They fail to close the nearly 20% gap to oracle performance, persisting in unproductive loops. We release the dataset and evaluation harness to help facilitate the transition from brute-force retrieval to calibrated, efficient reasoning.
APA
Borchmann, Ł., Van Landeghem, J., Turski, M., Padarha, S., Kearns, R.O., Mahdi, A., Rogge, N., Fourrier, C., Han, S., Yao, H., Llabrés, A., Xu, Y., Karatzas, D., Zhang, H. & Datta, A.. (2026). Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:9178-9241 Available from https://proceedings.mlr.press/v306/borchmann26a.html.

Related Material