Long term sequential decision making under risk

Mohammad Mirzanejad, Nadjet Bourdache, Abdel-illah Mouaddib
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:4543-4560, 2026.

Abstract

We study finite-horizon {MDP} planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break {Bellman} optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank–quantile surrogate via exact {DP} (Dynamic Programming), evaluates candidate policies exactly by {DP} over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper–lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-mirzanejad26a, title = {Long term sequential decision making under risk}, author = {Mirzanejad, Mohammad and Bourdache, Nadjet and Mouaddib, Abdel-illah}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {4543--4560}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/mirzanejad26a/mirzanejad26a.pdf}, url = {https://proceedings.mlr.press/v337/mirzanejad26a.html}, abstract = {We study finite-horizon {MDP} planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break {Bellman} optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank–quantile surrogate via exact {DP} (Dynamic Programming), evaluates candidate policies exactly by {DP} over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper–lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.} }
Endnote
%0 Conference Paper %T Long term sequential decision making under risk %A Mohammad Mirzanejad %A Nadjet Bourdache %A Abdel-illah Mouaddib %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-mirzanejad26a %I PMLR %P 4543--4560 %U https://proceedings.mlr.press/v337/mirzanejad26a.html %V 337 %X We study finite-horizon {MDP} planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break {Bellman} optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank–quantile surrogate via exact {DP} (Dynamic Programming), evaluates candidate policies exactly by {DP} over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper–lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.
APA
Mirzanejad, M., Bourdache, N. & Mouaddib, A.. (2026). Long term sequential decision making under risk. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:4543-4560 Available from https://proceedings.mlr.press/v337/mirzanejad26a.html.

Related Material