DRL-ORA: Distributional Reinforcement Learning with Online Epistemic Risk Adaptation

Yupeng Wu, Wenyun Li, Wenjie Huang, Chin Pang Ho
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7484-7503, 2026.

Abstract

One of the main challenges in reinforcement learning ({RL}) is that the agent has to make decisions that would influence the future performance without having complete knowledge of the environment. Dynamically adjusting the level of epistemic risk during the learning process can help to achieve reliable policies in safety-critical settings with better efficiency. In this work, we propose a new framework, Distributional {RL} with Online Epistemic Risk Adaptation ({DRL-ORA}). This framework quantifies both epistemic and implicit aleatory uncertainties in a unified manner and dynamically adjusts the epistemic risk levels by solving a total variation minimization problem online. The framework generalizes the existing variants of risk adaptation approaches with better explainability and flexibility. The selection of risk levels is performed efficiently via a Follow-The-Leader-type algorithm. We show that {DRL-ORA} outperforms existing methods that rely on fixed risk levels or manually designed risk level adaptation in multiple classes of tasks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-wu26e, title = {{DRL-ORA}: Distributional Reinforcement Learning with Online Epistemic Risk Adaptation}, author = {Wu, Yupeng and Li, Wenyun and Huang, Wenjie and Ho, Chin Pang}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7484--7503}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/wu26e/wu26e.pdf}, url = {https://proceedings.mlr.press/v337/wu26e.html}, abstract = {One of the main challenges in reinforcement learning ({RL}) is that the agent has to make decisions that would influence the future performance without having complete knowledge of the environment. Dynamically adjusting the level of epistemic risk during the learning process can help to achieve reliable policies in safety-critical settings with better efficiency. In this work, we propose a new framework, Distributional {RL} with Online Epistemic Risk Adaptation ({DRL-ORA}). This framework quantifies both epistemic and implicit aleatory uncertainties in a unified manner and dynamically adjusts the epistemic risk levels by solving a total variation minimization problem online. The framework generalizes the existing variants of risk adaptation approaches with better explainability and flexibility. The selection of risk levels is performed efficiently via a Follow-The-Leader-type algorithm. We show that {DRL-ORA} outperforms existing methods that rely on fixed risk levels or manually designed risk level adaptation in multiple classes of tasks.} }
Endnote
%0 Conference Paper %T DRL-ORA: Distributional Reinforcement Learning with Online Epistemic Risk Adaptation %A Yupeng Wu %A Wenyun Li %A Wenjie Huang %A Chin Pang Ho %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-wu26e %I PMLR %P 7484--7503 %U https://proceedings.mlr.press/v337/wu26e.html %V 337 %X One of the main challenges in reinforcement learning ({RL}) is that the agent has to make decisions that would influence the future performance without having complete knowledge of the environment. Dynamically adjusting the level of epistemic risk during the learning process can help to achieve reliable policies in safety-critical settings with better efficiency. In this work, we propose a new framework, Distributional {RL} with Online Epistemic Risk Adaptation ({DRL-ORA}). This framework quantifies both epistemic and implicit aleatory uncertainties in a unified manner and dynamically adjusts the epistemic risk levels by solving a total variation minimization problem online. The framework generalizes the existing variants of risk adaptation approaches with better explainability and flexibility. The selection of risk levels is performed efficiently via a Follow-The-Leader-type algorithm. We show that {DRL-ORA} outperforms existing methods that rely on fixed risk levels or manually designed risk level adaptation in multiple classes of tasks.
APA
Wu, Y., Li, W., Huang, W. & Ho, C.P.. (2026). DRL-ORA: Distributional Reinforcement Learning with Online Epistemic Risk Adaptation. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7484-7503 Available from https://proceedings.mlr.press/v337/wu26e.html.

Related Material