When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs

Jose Efraim Aguilar Escamilla, Haoyang Hong, Jiawei Li, Haoyu Zhao, Xuezhou Zhang, Sanghyun Hong, Huazheng Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:28237-28284, 2026.

Abstract

We study reward poisoning attacks in reinforcement learning (RL), where an adversary manipulates rewards under a limited budget to induce a target agent to learn a policy aligned with the attacker’s objectives. Most prior work focuses on constructing successful attacks, providing sufficient conditions under which poisoning is effective, while offering limited understanding of when such targeted attacks are fundamentally infeasible. In this paper, we provide the first characterization of reward-poisoning attackability in linear MDPs, establishing both necessary and sufficient conditions for whether a target policy can be induced within a bounded attack budget. This draws a clear boundary between the vulnerable RL instances and intrinsically robust ones, which cannot be attacked without high costs even when the learner uses standard, non-robust RL algorithms. We further demonstrate our framework beyond synthetic linear MDPs by approximating deep RL environments as linear MDPs. We show that our theoretical framework effectively distinguishes vulnerability, demonstrating how our theoretical predictions have practical significance.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-escamilla26a, title = {When Can You Poison Rewards? {A} Tight Characterization of Reward Poisoning in Linear {MDP}s}, author = {Escamilla, Jose Efraim Aguilar and Hong, Haoyang and Li, Jiawei and Zhao, Haoyu and Zhang, Xuezhou and Hong, Sanghyun and Wang, Huazheng}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {28237--28284}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/escamilla26a/escamilla26a.pdf}, url = {https://proceedings.mlr.press/v306/escamilla26a.html}, abstract = {We study reward poisoning attacks in reinforcement learning (RL), where an adversary manipulates rewards under a limited budget to induce a target agent to learn a policy aligned with the attacker’s objectives. Most prior work focuses on constructing successful attacks, providing sufficient conditions under which poisoning is effective, while offering limited understanding of when such targeted attacks are fundamentally infeasible. In this paper, we provide the first characterization of reward-poisoning attackability in linear MDPs, establishing both necessary and sufficient conditions for whether a target policy can be induced within a bounded attack budget. This draws a clear boundary between the vulnerable RL instances and intrinsically robust ones, which cannot be attacked without high costs even when the learner uses standard, non-robust RL algorithms. We further demonstrate our framework beyond synthetic linear MDPs by approximating deep RL environments as linear MDPs. We show that our theoretical framework effectively distinguishes vulnerability, demonstrating how our theoretical predictions have practical significance.} }
Endnote
%0 Conference Paper %T When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs %A Jose Efraim Aguilar Escamilla %A Haoyang Hong %A Jiawei Li %A Haoyu Zhao %A Xuezhou Zhang %A Sanghyun Hong %A Huazheng Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-escamilla26a %I PMLR %P 28237--28284 %U https://proceedings.mlr.press/v306/escamilla26a.html %V 306 %X We study reward poisoning attacks in reinforcement learning (RL), where an adversary manipulates rewards under a limited budget to induce a target agent to learn a policy aligned with the attacker’s objectives. Most prior work focuses on constructing successful attacks, providing sufficient conditions under which poisoning is effective, while offering limited understanding of when such targeted attacks are fundamentally infeasible. In this paper, we provide the first characterization of reward-poisoning attackability in linear MDPs, establishing both necessary and sufficient conditions for whether a target policy can be induced within a bounded attack budget. This draws a clear boundary between the vulnerable RL instances and intrinsically robust ones, which cannot be attacked without high costs even when the learner uses standard, non-robust RL algorithms. We further demonstrate our framework beyond synthetic linear MDPs by approximating deep RL environments as linear MDPs. We show that our theoretical framework effectively distinguishes vulnerability, demonstrating how our theoretical predictions have practical significance.
APA
Escamilla, J.E.A., Hong, H., Li, J., Zhao, H., Zhang, X., Hong, S. & Wang, H.. (2026). When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:28237-28284 Available from https://proceedings.mlr.press/v306/escamilla26a.html.

Related Material