Efficient Q-Learning and Actor–Critic Methods for Robust Average-Reward Reinforcement Learning.

Yang Xu, Swetha Ganesh, Vaneet Aggarwal
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7613-7650, 2026.

Abstract

We study model-free methods for distributionally robust infinite-horizon average-reward {Markov} decision processes ({MDPs}). We present non-asymptotic convergence analyses of $Q$-learning and actor–critic algorithms for robust average-reward {MDPs} under contamination, total-variation distance, and {Wasserstein} uncertainty sets. A key ingredient of our analysis is showing that the optimal robust {Bellman} operator is a strict contraction with respect to a carefully designed semi-norm. This property enables a stochastic approximation update that learns the optimal robust $Q$-function using $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We also establish robust {TD} convergence bounds whose constants are uniform over all stationary policies, yielding an efficient data-driven routine for robust critic estimation. Building on this, we introduce an actor–critic algorithm that learns an $\epsilon$-optimal robust policy within $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We provide numerical simulations to illustrate the qualitative behavior of the proposed algorithms. Our results contribute to the theoretical foundations of robust planning under model misspecification, and to model-free approaches for building robust long-run policies directly from simulation data.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-xu26d, title = {Efficient {Q-Learning} and Actor–Critic Methods for Robust Average-Reward Reinforcement Learning.}, author = {Xu, Yang and Ganesh, Swetha and Aggarwal, Vaneet}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7613--7650}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/xu26d/xu26d.pdf}, url = {https://proceedings.mlr.press/v337/xu26d.html}, abstract = {We study model-free methods for distributionally robust infinite-horizon average-reward {Markov} decision processes ({MDPs}). We present non-asymptotic convergence analyses of $Q$-learning and actor–critic algorithms for robust average-reward {MDPs} under contamination, total-variation distance, and {Wasserstein} uncertainty sets. A key ingredient of our analysis is showing that the optimal robust {Bellman} operator is a strict contraction with respect to a carefully designed semi-norm. This property enables a stochastic approximation update that learns the optimal robust $Q$-function using $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We also establish robust {TD} convergence bounds whose constants are uniform over all stationary policies, yielding an efficient data-driven routine for robust critic estimation. Building on this, we introduce an actor–critic algorithm that learns an $\epsilon$-optimal robust policy within $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We provide numerical simulations to illustrate the qualitative behavior of the proposed algorithms. Our results contribute to the theoretical foundations of robust planning under model misspecification, and to model-free approaches for building robust long-run policies directly from simulation data.} }
Endnote
%0 Conference Paper %T Efficient Q-Learning and Actor–Critic Methods for Robust Average-Reward Reinforcement Learning. %A Yang Xu %A Swetha Ganesh %A Vaneet Aggarwal %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-xu26d %I PMLR %P 7613--7650 %U https://proceedings.mlr.press/v337/xu26d.html %V 337 %X We study model-free methods for distributionally robust infinite-horizon average-reward {Markov} decision processes ({MDPs}). We present non-asymptotic convergence analyses of $Q$-learning and actor–critic algorithms for robust average-reward {MDPs} under contamination, total-variation distance, and {Wasserstein} uncertainty sets. A key ingredient of our analysis is showing that the optimal robust {Bellman} operator is a strict contraction with respect to a carefully designed semi-norm. This property enables a stochastic approximation update that learns the optimal robust $Q$-function using $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We also establish robust {TD} convergence bounds whose constants are uniform over all stationary policies, yielding an efficient data-driven routine for robust critic estimation. Building on this, we introduce an actor–critic algorithm that learns an $\epsilon$-optimal robust policy within $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We provide numerical simulations to illustrate the qualitative behavior of the proposed algorithms. Our results contribute to the theoretical foundations of robust planning under model misspecification, and to model-free approaches for building robust long-run policies directly from simulation data.
APA
Xu, Y., Ganesh, S. & Aggarwal, V.. (2026). Efficient Q-Learning and Actor–Critic Methods for Robust Average-Reward Reinforcement Learning.. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7613-7650 Available from https://proceedings.mlr.press/v337/xu26d.html.

Related Material