Optimistic Reinforcement Learning with Quantile Objectives

Mohammad Alipour-Vaezi, Huaiyang Zhong, Kwok-leung Tsui, Sajad Khodadadian
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3601-3609, 2026.

Abstract

Reinforcement Learning (RL) has achieved tremendous success in recent years. However, the classical foundations of RL do not account for the risk sensitivity of the objective function, which is critical in various fields, including healthcare, finance, etc. A popular approach to incorporate risk sensitivity is to optimize a specific quantile of the cumulative reward distribution. In this paper, we develop UCB-QRL, an optimistic learning algorithm for the $\tau$-quantile objective in finite-horizon Markov decision processes (MDPs). UCB-QRL is an iterative algorithm in which, at each iteration, we first estimate the underlying transition probability and then optimize the quantile value function over a confidence ball around this estimate. Here, we show that UCB-QRL yields high-probability regret bounds $\mathcal O\left((2/\kappa)^HH\sqrt{SATH\log(2SATH/\delta)}\right)$ in the episodic setting with $S$ states, $A$ actions, $T$ episodes, and $H$ horizons. Here, $\kappa>0$ is a problem-dependent constant that captures the sensitivity of the underlying MDP’s quantile value.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-alipour-vaezi26a, title = { Optimistic Reinforcement Learning with Quantile Objectives }, author = {Alipour-Vaezi, Mohammad and Zhong, Huaiyang and Tsui, Kwok-leung and Khodadadian, Sajad}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3601--3609}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/alipour-vaezi26a/alipour-vaezi26a.pdf}, url = {https://proceedings.mlr.press/v300/alipour-vaezi26a.html}, abstract = { Reinforcement Learning (RL) has achieved tremendous success in recent years. However, the classical foundations of RL do not account for the risk sensitivity of the objective function, which is critical in various fields, including healthcare, finance, etc. A popular approach to incorporate risk sensitivity is to optimize a specific quantile of the cumulative reward distribution. In this paper, we develop UCB-QRL, an optimistic learning algorithm for the $\tau$-quantile objective in finite-horizon Markov decision processes (MDPs). UCB-QRL is an iterative algorithm in which, at each iteration, we first estimate the underlying transition probability and then optimize the quantile value function over a confidence ball around this estimate. Here, we show that UCB-QRL yields high-probability regret bounds $\mathcal O\left((2/\kappa)^HH\sqrt{SATH\log(2SATH/\delta)}\right)$ in the episodic setting with $S$ states, $A$ actions, $T$ episodes, and $H$ horizons. Here, $\kappa>0$ is a problem-dependent constant that captures the sensitivity of the underlying MDP’s quantile value. } }
Endnote
%0 Conference Paper %T Optimistic Reinforcement Learning with Quantile Objectives %A Mohammad Alipour-Vaezi %A Huaiyang Zhong %A Kwok-leung Tsui %A Sajad Khodadadian %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-alipour-vaezi26a %I PMLR %P 3601--3609 %U https://proceedings.mlr.press/v300/alipour-vaezi26a.html %V 300 %X Reinforcement Learning (RL) has achieved tremendous success in recent years. However, the classical foundations of RL do not account for the risk sensitivity of the objective function, which is critical in various fields, including healthcare, finance, etc. A popular approach to incorporate risk sensitivity is to optimize a specific quantile of the cumulative reward distribution. In this paper, we develop UCB-QRL, an optimistic learning algorithm for the $\tau$-quantile objective in finite-horizon Markov decision processes (MDPs). UCB-QRL is an iterative algorithm in which, at each iteration, we first estimate the underlying transition probability and then optimize the quantile value function over a confidence ball around this estimate. Here, we show that UCB-QRL yields high-probability regret bounds $\mathcal O\left((2/\kappa)^HH\sqrt{SATH\log(2SATH/\delta)}\right)$ in the episodic setting with $S$ states, $A$ actions, $T$ episodes, and $H$ horizons. Here, $\kappa>0$ is a problem-dependent constant that captures the sensitivity of the underlying MDP’s quantile value.
APA
Alipour-Vaezi, M., Zhong, H., Tsui, K. & Khodadadian, S.. (2026). Optimistic Reinforcement Learning with Quantile Objectives . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3601-3609 Available from https://proceedings.mlr.press/v300/alipour-vaezi26a.html.

Related Material