Performative Policy Gradient: Optimality in Performative Reinforcement Learning

Debabrota Basu, Udvas Das, Brahim Driss, Uddalak Mukherjee
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:6939-6983, 2026.

Abstract

Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) methods ignore. While designing optimal algorithms in this performative setting has recently been studied in supervised learning, the RL counterpart remains under-explored. In this paper, we prove the performative counterparts of the performance difference lemma and the policy gradient theorem in RL, and introduce the Performative Policy Gradient algorithm PePG. PePG is the first policy gradient algorithm designed to account for performativity in RL. Under softmax parametrisation, and also with and without entropy regularisation, we prove that PePG converges to performatively optimal policies, i.e. policies that remain optimal under the distribution shifts induced by themselves. Thus, PePG significantly extends the prior works in Performative RL that achieves performative stability but not optimality. Our empirical analysis on standard performative RL environments validate that PePG outperforms the existing performative RL algorithms aiming for stability.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-basu26a, title = {Performative Policy Gradient: Optimality in Performative Reinforcement Learning}, author = {Basu, Debabrota and Das, Udvas and Driss, Brahim and Mukherjee, Uddalak}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {6939--6983}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/basu26a/basu26a.pdf}, url = {https://proceedings.mlr.press/v306/basu26a.html}, abstract = {Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) methods ignore. While designing optimal algorithms in this performative setting has recently been studied in supervised learning, the RL counterpart remains under-explored. In this paper, we prove the performative counterparts of the performance difference lemma and the policy gradient theorem in RL, and introduce the Performative Policy Gradient algorithm PePG. PePG is the first policy gradient algorithm designed to account for performativity in RL. Under softmax parametrisation, and also with and without entropy regularisation, we prove that PePG converges to performatively optimal policies, i.e. policies that remain optimal under the distribution shifts induced by themselves. Thus, PePG significantly extends the prior works in Performative RL that achieves performative stability but not optimality. Our empirical analysis on standard performative RL environments validate that PePG outperforms the existing performative RL algorithms aiming for stability.} }
Endnote
%0 Conference Paper %T Performative Policy Gradient: Optimality in Performative Reinforcement Learning %A Debabrota Basu %A Udvas Das %A Brahim Driss %A Uddalak Mukherjee %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-basu26a %I PMLR %P 6939--6983 %U https://proceedings.mlr.press/v306/basu26a.html %V 306 %X Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) methods ignore. While designing optimal algorithms in this performative setting has recently been studied in supervised learning, the RL counterpart remains under-explored. In this paper, we prove the performative counterparts of the performance difference lemma and the policy gradient theorem in RL, and introduce the Performative Policy Gradient algorithm PePG. PePG is the first policy gradient algorithm designed to account for performativity in RL. Under softmax parametrisation, and also with and without entropy regularisation, we prove that PePG converges to performatively optimal policies, i.e. policies that remain optimal under the distribution shifts induced by themselves. Thus, PePG significantly extends the prior works in Performative RL that achieves performative stability but not optimality. Our empirical analysis on standard performative RL environments validate that PePG outperforms the existing performative RL algorithms aiming for stability.
APA
Basu, D., Das, U., Driss, B. & Mukherjee, U.. (2026). Performative Policy Gradient: Optimality in Performative Reinforcement Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:6939-6983 Available from https://proceedings.mlr.press/v306/basu26a.html.

Related Material