AgentSuite: Toward More Reliable Agent Evaluation with a Component-Based Benchmark Auditing Pipeline

Hyewon Suh, Binfei Ji, Seojune Lee, Rishi Khare, Basit Khan, Hyunjun Kim, Tianyi Zhang, Venkat Krishna Srinivasan, Peter Belcak, Shizhe Diao, Pavlo Molchanov, Yingyan Celine Lin, Zhen Dong
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:116300-116319, 2026.

Abstract

Reliable evaluation of large language model (LLM) agents depends critically on benchmark validity, yet modern agent benchmarks often contain hidden flaws arising from interactions among user instructions, environments, tools, ground-truth trajectories, and evaluation protocols. These flaws confound model errors with benchmark artifacts and undermine leaderboard-based comparisons. We propose COBA (COmponent-based Benchmark Auditing), an automated pipeline for diagnosing and filtering validity issues in agent benchmarks. COBA decomposes agent tasks into four standardized components—User, Environment, Ground Truth, and Evaluation—and operationalizes a component-level issue taxonomy using hybrid rule-based detectors and taxonomy-guided LLM evaluation. Across six widely used agent benchmarks, COBA achieves strong alignment with expert judgments, with F1 scores between 0.791 and 0.874. It complements manual verification of $\tau^2$-bench by identifying issues missed due to benchmark complexity, and generalizes to previously unseen benchmarks with minimal adaptation. Our analysis shows that benchmark flaws are widespread and materially affect evaluation outcomes, demonstrating that component-based auditing provides a scalable foundation for more reliable and interpretable agent evaluation. We release AgentSuite, a unified benchmark-running platform that includes the COBA auditing pipeline and audited benchmark variants: https://github.com/Agent-Suite/AgentSuite.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-suh26a, title = {{A}gent{S}uite: Toward More Reliable Agent Evaluation with a Component-Based Benchmark Auditing Pipeline}, author = {Suh, Hyewon and Ji, Binfei and Lee, Seojune and Khare, Rishi and Khan, Basit and Kim, Hyunjun and Zhang, Tianyi and Srinivasan, Venkat Krishna and Belcak, Peter and Diao, Shizhe and Molchanov, Pavlo and Lin, Yingyan Celine and Dong, Zhen}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {116300--116319}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/suh26a/suh26a.pdf}, url = {https://proceedings.mlr.press/v306/suh26a.html}, abstract = {Reliable evaluation of large language model (LLM) agents depends critically on benchmark validity, yet modern agent benchmarks often contain hidden flaws arising from interactions among user instructions, environments, tools, ground-truth trajectories, and evaluation protocols. These flaws confound model errors with benchmark artifacts and undermine leaderboard-based comparisons. We propose COBA (COmponent-based Benchmark Auditing), an automated pipeline for diagnosing and filtering validity issues in agent benchmarks. COBA decomposes agent tasks into four standardized components—User, Environment, Ground Truth, and Evaluation—and operationalizes a component-level issue taxonomy using hybrid rule-based detectors and taxonomy-guided LLM evaluation. Across six widely used agent benchmarks, COBA achieves strong alignment with expert judgments, with F1 scores between 0.791 and 0.874. It complements manual verification of $\tau^2$-bench by identifying issues missed due to benchmark complexity, and generalizes to previously unseen benchmarks with minimal adaptation. Our analysis shows that benchmark flaws are widespread and materially affect evaluation outcomes, demonstrating that component-based auditing provides a scalable foundation for more reliable and interpretable agent evaluation. We release AgentSuite, a unified benchmark-running platform that includes the COBA auditing pipeline and audited benchmark variants: https://github.com/Agent-Suite/AgentSuite.} }
Endnote
%0 Conference Paper %T AgentSuite: Toward More Reliable Agent Evaluation with a Component-Based Benchmark Auditing Pipeline %A Hyewon Suh %A Binfei Ji %A Seojune Lee %A Rishi Khare %A Basit Khan %A Hyunjun Kim %A Tianyi Zhang %A Venkat Krishna Srinivasan %A Peter Belcak %A Shizhe Diao %A Pavlo Molchanov %A Yingyan Celine Lin %A Zhen Dong %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-suh26a %I PMLR %P 116300--116319 %U https://proceedings.mlr.press/v306/suh26a.html %V 306 %X Reliable evaluation of large language model (LLM) agents depends critically on benchmark validity, yet modern agent benchmarks often contain hidden flaws arising from interactions among user instructions, environments, tools, ground-truth trajectories, and evaluation protocols. These flaws confound model errors with benchmark artifacts and undermine leaderboard-based comparisons. We propose COBA (COmponent-based Benchmark Auditing), an automated pipeline for diagnosing and filtering validity issues in agent benchmarks. COBA decomposes agent tasks into four standardized components—User, Environment, Ground Truth, and Evaluation—and operationalizes a component-level issue taxonomy using hybrid rule-based detectors and taxonomy-guided LLM evaluation. Across six widely used agent benchmarks, COBA achieves strong alignment with expert judgments, with F1 scores between 0.791 and 0.874. It complements manual verification of $\tau^2$-bench by identifying issues missed due to benchmark complexity, and generalizes to previously unseen benchmarks with minimal adaptation. Our analysis shows that benchmark flaws are widespread and materially affect evaluation outcomes, demonstrating that component-based auditing provides a scalable foundation for more reliable and interpretable agent evaluation. We release AgentSuite, a unified benchmark-running platform that includes the COBA auditing pipeline and audited benchmark variants: https://github.com/Agent-Suite/AgentSuite.
APA
Suh, H., Ji, B., Lee, S., Khare, R., Khan, B., Kim, H., Zhang, T., Srinivasan, V.K., Belcak, P., Diao, S., Molchanov, P., Lin, Y.C. & Dong, Z.. (2026). AgentSuite: Toward More Reliable Agent Evaluation with a Component-Based Benchmark Auditing Pipeline. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:116300-116319 Available from https://proceedings.mlr.press/v306/suh26a.html.

Related Material