Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization

Anirudh Satheesh, Pankaj Kumar Barman, Washim Uddin Mondal, Vaneet Aggarwal
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:5997-6025, 2026.

Abstract

We study infinite-horizon Constrained {Markov} Decision Processes (CMDPs) with general policy parameterizations and multi-layer neural network critics. Existing theoretical analyses for constrained reinforcement learning largely rely on tabular policies or linear critics, which limits their applicability to high-dimensional and continuous control problems. We propose a primal–dual natural actor–critic algorithm that integrates neural critic estimation with natural policy gradient updates and leverages Neural Tangent Kernel ({NTK}) theory to control function-approximation error under Markovian sampling, without requiring access to mixing-time oracles. We establish global convergence and cumulative constraint violation rates of $\tilde{\mathcal{O}}(T^{-1/4})$ up to approximation errors induced by the policy and critic classes. Our results provide the first such guarantees for CMDPs with general policies and multi-layer neural critics, substantially extending the theoretical foundations of actor–critic methods beyond the linear-critic regime.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-satheesh26a, title = {Global Convergence of Average Reward Constrained {MDPs} with Neural Critic and General Policy Parameterization}, author = {Satheesh, Anirudh and Barman, Pankaj Kumar and Mondal, Washim Uddin and Aggarwal, Vaneet}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {5997--6025}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/satheesh26a/satheesh26a.pdf}, url = {https://proceedings.mlr.press/v337/satheesh26a.html}, abstract = {We study infinite-horizon Constrained {Markov} Decision Processes (CMDPs) with general policy parameterizations and multi-layer neural network critics. Existing theoretical analyses for constrained reinforcement learning largely rely on tabular policies or linear critics, which limits their applicability to high-dimensional and continuous control problems. We propose a primal–dual natural actor–critic algorithm that integrates neural critic estimation with natural policy gradient updates and leverages Neural Tangent Kernel ({NTK}) theory to control function-approximation error under Markovian sampling, without requiring access to mixing-time oracles. We establish global convergence and cumulative constraint violation rates of $\tilde{\mathcal{O}}(T^{-1/4})$ up to approximation errors induced by the policy and critic classes. Our results provide the first such guarantees for CMDPs with general policies and multi-layer neural critics, substantially extending the theoretical foundations of actor–critic methods beyond the linear-critic regime.} }
Endnote
%0 Conference Paper %T Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization %A Anirudh Satheesh %A Pankaj Kumar Barman %A Washim Uddin Mondal %A Vaneet Aggarwal %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-satheesh26a %I PMLR %P 5997--6025 %U https://proceedings.mlr.press/v337/satheesh26a.html %V 337 %X We study infinite-horizon Constrained {Markov} Decision Processes (CMDPs) with general policy parameterizations and multi-layer neural network critics. Existing theoretical analyses for constrained reinforcement learning largely rely on tabular policies or linear critics, which limits their applicability to high-dimensional and continuous control problems. We propose a primal–dual natural actor–critic algorithm that integrates neural critic estimation with natural policy gradient updates and leverages Neural Tangent Kernel ({NTK}) theory to control function-approximation error under Markovian sampling, without requiring access to mixing-time oracles. We establish global convergence and cumulative constraint violation rates of $\tilde{\mathcal{O}}(T^{-1/4})$ up to approximation errors induced by the policy and critic classes. Our results provide the first such guarantees for CMDPs with general policies and multi-layer neural critics, substantially extending the theoretical foundations of actor–critic methods beyond the linear-critic regime.
APA
Satheesh, A., Barman, P.K., Mondal, W.U. & Aggarwal, V.. (2026). Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:5997-6025 Available from https://proceedings.mlr.press/v337/satheesh26a.html.

Related Material