AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks

Tanqiu Jiang, Yuhui Wang, Jiacheng Liang, Ting Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:53162-53180, 2026.

Abstract

LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user–agent–environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. Currently, AgentLAB supports five novel attack types including intent hijacking, tool chaining, task injection, objective drifting, and memory poisoning, spanning 28 realistic agentic environments, and 644 security test cases. Leveraging AgentLAB, we evaluate representative LLM agents and find that they remain highly susceptible to long-horizon attacks; moreover, defenses designed for single-turn interactions fail to reliably mitigate long-horizon threats. We anticipate that AgentLAB will serve as a valuable benchmark for tracking progress on securing LLM agents in practical settings. The code for AgentLAB is available at: https://tanqiujiang.github.io/AgentLAB_main.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-jiang26as, title = {{A}gent{LAB}: Benchmarking {LLM} Agents against Long-Horizon Attacks}, author = {Jiang, Tanqiu and Wang, Yuhui and Liang, Jiacheng and Wang, Ting}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {53162--53180}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/jiang26as/jiang26as.pdf}, url = {https://proceedings.mlr.press/v306/jiang26as.html}, abstract = {LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user–agent–environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. Currently, AgentLAB supports five novel attack types including intent hijacking, tool chaining, task injection, objective drifting, and memory poisoning, spanning 28 realistic agentic environments, and 644 security test cases. Leveraging AgentLAB, we evaluate representative LLM agents and find that they remain highly susceptible to long-horizon attacks; moreover, defenses designed for single-turn interactions fail to reliably mitigate long-horizon threats. We anticipate that AgentLAB will serve as a valuable benchmark for tracking progress on securing LLM agents in practical settings. The code for AgentLAB is available at: https://tanqiujiang.github.io/AgentLAB_main.} }
Endnote
%0 Conference Paper %T AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks %A Tanqiu Jiang %A Yuhui Wang %A Jiacheng Liang %A Ting Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-jiang26as %I PMLR %P 53162--53180 %U https://proceedings.mlr.press/v306/jiang26as.html %V 306 %X LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user–agent–environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. Currently, AgentLAB supports five novel attack types including intent hijacking, tool chaining, task injection, objective drifting, and memory poisoning, spanning 28 realistic agentic environments, and 644 security test cases. Leveraging AgentLAB, we evaluate representative LLM agents and find that they remain highly susceptible to long-horizon attacks; moreover, defenses designed for single-turn interactions fail to reliably mitigate long-horizon threats. We anticipate that AgentLAB will serve as a valuable benchmark for tracking progress on securing LLM agents in practical settings. The code for AgentLAB is available at: https://tanqiujiang.github.io/AgentLAB_main.
APA
Jiang, T., Wang, Y., Liang, J. & Wang, T.. (2026). AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:53162-53180 Available from https://proceedings.mlr.press/v306/jiang26as.html.

Related Material