[edit]
Efficient Q-Learning and Actor–Critic Methods for Robust Average-Reward Reinforcement Learning.
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7613-7650, 2026.
Abstract
We study model-free methods for distributionally robust infinite-horizon average-reward {Markov} decision processes ({MDPs}). We present non-asymptotic convergence analyses of $Q$-learning and actor–critic algorithms for robust average-reward {MDPs} under contamination, total-variation distance, and {Wasserstein} uncertainty sets. A key ingredient of our analysis is showing that the optimal robust {Bellman} operator is a strict contraction with respect to a carefully designed semi-norm. This property enables a stochastic approximation update that learns the optimal robust $Q$-function using $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We also establish robust {TD} convergence bounds whose constants are uniform over all stationary policies, yielding an efficient data-driven routine for robust critic estimation. Building on this, we introduce an actor–critic algorithm that learns an $\epsilon$-optimal robust policy within $\tilde{\mathcal{O}}(\epsilon^{-2})$ samples. We provide numerical simulations to illustrate the qualitative behavior of the proposed algorithms. Our results contribute to the theoretical foundations of robust planning under model misspecification, and to model-free approaches for building robust long-run policies directly from simulation data.