[edit]
Stochastic Dominance Driven First-Order Policy Optimization for Multi-Objective Reinforcement Learning
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2828-2876, 2026.
Abstract
We study multi-objective reinforcement learning with stochastic, vector-valued returns and propose a policy improvement principle based on multivariate $k$th-order stochastic dominance in the lower-orthant sense. Rather than scalarizing objectives or comparing only expected values, we compare entire return distributions using $k$th-order integrated cumulative distribution functions (cdfs). This yields a max-violation objective that measures the worst-case dominance gap of a candidate policy relative to the incumbent policy. Under some regularity assumptions, we show that the Orthant Dominance Policy Optimization (ORDO) algorithm we introduce drives this objective to a point that certifies $\epsilon$-almost non-dominatedness, meaning that no policy in the class can improve the incumbent by more than $\epsilon$. We also propose smoothing approaches for the hinge structure inherent in integrated cdfs, which improve stability and enable stable stochastic gradients in optimization.