Parametric Return Density Estimation for Reinforcement Learning

Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, Toshiyuki Tanaka
Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, PMLR R8:375-382, 2010.

Abstract

Most conventional Reinforcement Learning (RL) algorithms aim to optimize decision- making rules in terms of the expected re- turns. However, especially for risk man- agement purposes, other risk-sensitive crite- ria such as the value-at-risk or the expected shortfall are sometimes preferred in real ap- plications. Here, we describe a parametric method for estimating density of the returns, which allows us to handle various criteria in a unified manner. We first extend the Bellman equation for the conditional expected return to cover a conditional probability density of the returns. Then we derive an extension of the TD-learning algorithm for estimating the return densities in an unknown environment. As test instances, several parametric density estimation algorithms are presented for the Gaussian, Laplace, and skewed Laplace dis- tributions. We show that these algorithms lead to risk-sensitive as well as robust RL paradigms through numerical experiments.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR8-morimura10a, title = {Parametric Return Density Estimation for Reinforcement Learning}, author = {Morimura, Tetsuro and Sugiyama, Masashi and Kashima, Hisashi and Hachiya, Hirotaka and Tanaka, Toshiyuki}, booktitle = {Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence}, pages = {375--382}, year = {2010}, editor = {Grünwald, Peter and Spirtes, Peter}, volume = {R8}, series = {Proceedings of Machine Learning Research}, month = {08--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r8/main/assets/morimura10a/morimura10a.pdf}, url = {https://proceedings.mlr.press/r8/morimura10a.html}, abstract = {Most conventional Reinforcement Learning (RL) algorithms aim to optimize decision- making rules in terms of the expected re- turns. However, especially for risk man- agement purposes, other risk-sensitive crite- ria such as the value-at-risk or the expected shortfall are sometimes preferred in real ap- plications. Here, we describe a parametric method for estimating density of the returns, which allows us to handle various criteria in a unified manner. We first extend the Bellman equation for the conditional expected return to cover a conditional probability density of the returns. Then we derive an extension of the TD-learning algorithm for estimating the return densities in an unknown environment. As test instances, several parametric density estimation algorithms are presented for the Gaussian, Laplace, and skewed Laplace dis- tributions. We show that these algorithms lead to risk-sensitive as well as robust RL paradigms through numerical experiments.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Parametric Return Density Estimation for Reinforcement Learning %A Tetsuro Morimura %A Masashi Sugiyama %A Hisashi Kashima %A Hirotaka Hachiya %A Toshiyuki Tanaka %B Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2010 %E Peter Grünwald %E Peter Spirtes %F pmlr-vR8-morimura10a %I PMLR %P 375--382 %U https://proceedings.mlr.press/r8/morimura10a.html %V R8 %X Most conventional Reinforcement Learning (RL) algorithms aim to optimize decision- making rules in terms of the expected re- turns. However, especially for risk man- agement purposes, other risk-sensitive crite- ria such as the value-at-risk or the expected shortfall are sometimes preferred in real ap- plications. Here, we describe a parametric method for estimating density of the returns, which allows us to handle various criteria in a unified manner. We first extend the Bellman equation for the conditional expected return to cover a conditional probability density of the returns. Then we derive an extension of the TD-learning algorithm for estimating the return densities in an unknown environment. As test instances, several parametric density estimation algorithms are presented for the Gaussian, Laplace, and skewed Laplace dis- tributions. We show that these algorithms lead to risk-sensitive as well as robust RL paradigms through numerical experiments. %Z Reissued by PMLR on 04 October 2026.
APA
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H. & Tanaka, T.. (2010). Parametric Return Density Estimation for Reinforcement Learning. Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R8:375-382 Available from https://proceedings.mlr.press/r8/morimura10a.html. Reissued by PMLR on 04 October 2026.

Related Material