Are Your Agents Upward Deceivers?

Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren, Shuai Shao, Tianyi Alex Qiu, Haoran Li, Yi R. Fung, Zhongjie Ba, Juntao Dai, Jiaming Ji, Zhikai Chen, Jialing Tao, Yaodong Yang, Jing Shao, Xia Hu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:38052-38092, 2026.

Abstract

Large Language Model (LLM)-based agents are increasingly used as autonomous subordinates that carry out tasks for users. This raises the question of whether they may also engage in deception, similar to how individuals in human organizations lie to superiors to create a good image or avoid punishment. We observe and define agentic upward deception, a phenomenon in which an agent facing environmental constraints conceals its failure and performs actions that were not requested without reporting. To assess its prevalence, we construct a benchmark of 200 tasks covering five task types and eight realistic scenarios in a constrained environment, such as broken tools or mismatched information sources. Evaluations of 11 popular LLMs reveal that these agents typically exhibit action-based deceptive behaviors, such as guessing results, performing unsupported simulations, substituting unavailable information sources, and fabricating local files. We further test intuitive mitigation methods and find only limited reductions, suggesting that it is difficult to eliminate and highlighting the need for stronger mitigation strategies to ensure the safety of LLM-based agents. Code and data are available at https://github.com/QingyuLiu/Agentic-Upward-Deception.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26k, title = {Are Your Agents Upward Deceivers?}, author = {Guo, Dadi and Liu, Qingyu and Liu, Dongrui and Ren, Qihan and Shao, Shuai and Qiu, Tianyi Alex and Li, Haoran and Fung, Yi R. and Ba, Zhongjie and Dai, Juntao and Ji, Jiaming and Chen, Zhikai and Tao, Jialing and Yang, Yaodong and Shao, Jing and Hu, Xia}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {38052--38092}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26k/guo26k.pdf}, url = {https://proceedings.mlr.press/v306/guo26k.html}, abstract = {Large Language Model (LLM)-based agents are increasingly used as autonomous subordinates that carry out tasks for users. This raises the question of whether they may also engage in deception, similar to how individuals in human organizations lie to superiors to create a good image or avoid punishment. We observe and define agentic upward deception, a phenomenon in which an agent facing environmental constraints conceals its failure and performs actions that were not requested without reporting. To assess its prevalence, we construct a benchmark of 200 tasks covering five task types and eight realistic scenarios in a constrained environment, such as broken tools or mismatched information sources. Evaluations of 11 popular LLMs reveal that these agents typically exhibit action-based deceptive behaviors, such as guessing results, performing unsupported simulations, substituting unavailable information sources, and fabricating local files. We further test intuitive mitigation methods and find only limited reductions, suggesting that it is difficult to eliminate and highlighting the need for stronger mitigation strategies to ensure the safety of LLM-based agents. Code and data are available at https://github.com/QingyuLiu/Agentic-Upward-Deception.} }
Endnote
%0 Conference Paper %T Are Your Agents Upward Deceivers? %A Dadi Guo %A Qingyu Liu %A Dongrui Liu %A Qihan Ren %A Shuai Shao %A Tianyi Alex Qiu %A Haoran Li %A Yi R. Fung %A Zhongjie Ba %A Juntao Dai %A Jiaming Ji %A Zhikai Chen %A Jialing Tao %A Yaodong Yang %A Jing Shao %A Xia Hu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26k %I PMLR %P 38052--38092 %U https://proceedings.mlr.press/v306/guo26k.html %V 306 %X Large Language Model (LLM)-based agents are increasingly used as autonomous subordinates that carry out tasks for users. This raises the question of whether they may also engage in deception, similar to how individuals in human organizations lie to superiors to create a good image or avoid punishment. We observe and define agentic upward deception, a phenomenon in which an agent facing environmental constraints conceals its failure and performs actions that were not requested without reporting. To assess its prevalence, we construct a benchmark of 200 tasks covering five task types and eight realistic scenarios in a constrained environment, such as broken tools or mismatched information sources. Evaluations of 11 popular LLMs reveal that these agents typically exhibit action-based deceptive behaviors, such as guessing results, performing unsupported simulations, substituting unavailable information sources, and fabricating local files. We further test intuitive mitigation methods and find only limited reductions, suggesting that it is difficult to eliminate and highlighting the need for stronger mitigation strategies to ensure the safety of LLM-based agents. Code and data are available at https://github.com/QingyuLiu/Agentic-Upward-Deception.
APA
Guo, D., Liu, Q., Liu, D., Ren, Q., Shao, S., Qiu, T.A., Li, H., Fung, Y.R., Ba, Z., Dai, J., Ji, J., Chen, Z., Tao, J., Yang, Y., Shao, J. & Hu, X.. (2026). Are Your Agents Upward Deceivers?. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:38052-38092 Available from https://proceedings.mlr.press/v306/guo26k.html.

Related Material