Stochastic Dominance Driven First-Order Policy Optimization for Multi-Objective Reinforcement Learning

Ege Can Kaya, Kadierdan Kaheman, Jason M Cloud, Abolfazl Hashemi
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2828-2876, 2026.

Abstract

We study multi-objective reinforcement learning with stochastic, vector-valued returns and propose a policy improvement principle based on multivariate $k$th-order stochastic dominance in the lower-orthant sense. Rather than scalarizing objectives or comparing only expected values, we compare entire return distributions using $k$th-order integrated cumulative distribution functions (cdfs). This yields a max-violation objective that measures the worst-case dominance gap of a candidate policy relative to the incumbent policy. Under some regularity assumptions, we show that the Orthant Dominance Policy Optimization (ORDO) algorithm we introduce drives this objective to a point that certifies $\epsilon$-almost non-dominatedness, meaning that no policy in the class can improve the incumbent by more than $\epsilon$. We also propose smoothing approaches for the hinge structure inherent in integrated cdfs, which improve stability and enable stable stochastic gradients in optimization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kaya26a, title = {Stochastic Dominance Driven First-Order Policy Optimization for Multi-Objective Reinforcement Learning}, author = {Kaya, Ege Can and Kaheman, Kadierdan and Cloud, Jason M and Hashemi, Abolfazl}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2828--2876}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kaya26a/kaya26a.pdf}, url = {https://proceedings.mlr.press/v337/kaya26a.html}, abstract = {We study multi-objective reinforcement learning with stochastic, vector-valued returns and propose a policy improvement principle based on multivariate $k$th-order stochastic dominance in the lower-orthant sense. Rather than scalarizing objectives or comparing only expected values, we compare entire return distributions using $k$th-order integrated cumulative distribution functions (cdfs). This yields a max-violation objective that measures the worst-case dominance gap of a candidate policy relative to the incumbent policy. Under some regularity assumptions, we show that the Orthant Dominance Policy Optimization (ORDO) algorithm we introduce drives this objective to a point that certifies $\epsilon$-almost non-dominatedness, meaning that no policy in the class can improve the incumbent by more than $\epsilon$. We also propose smoothing approaches for the hinge structure inherent in integrated cdfs, which improve stability and enable stable stochastic gradients in optimization.} }
Endnote
%0 Conference Paper %T Stochastic Dominance Driven First-Order Policy Optimization for Multi-Objective Reinforcement Learning %A Ege Can Kaya %A Kadierdan Kaheman %A Jason M Cloud %A Abolfazl Hashemi %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kaya26a %I PMLR %P 2828--2876 %U https://proceedings.mlr.press/v337/kaya26a.html %V 337 %X We study multi-objective reinforcement learning with stochastic, vector-valued returns and propose a policy improvement principle based on multivariate $k$th-order stochastic dominance in the lower-orthant sense. Rather than scalarizing objectives or comparing only expected values, we compare entire return distributions using $k$th-order integrated cumulative distribution functions (cdfs). This yields a max-violation objective that measures the worst-case dominance gap of a candidate policy relative to the incumbent policy. Under some regularity assumptions, we show that the Orthant Dominance Policy Optimization (ORDO) algorithm we introduce drives this objective to a point that certifies $\epsilon$-almost non-dominatedness, meaning that no policy in the class can improve the incumbent by more than $\epsilon$. We also propose smoothing approaches for the hinge structure inherent in integrated cdfs, which improve stability and enable stable stochastic gradients in optimization.
APA
Kaya, E.C., Kaheman, K., Cloud, J.M. & Hashemi, A.. (2026). Stochastic Dominance Driven First-Order Policy Optimization for Multi-Objective Reinforcement Learning. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2828-2876 Available from https://proceedings.mlr.press/v337/kaya26a.html.

Related Material