Soft-Robust Actor-Critic Policy-Gradient

Esther Derman, Daniel J Mankowitz, Timothy A Mann, Shie Mannor
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:207-217, 2018.

Abstract

Robust Reinforcement Learning aims to derive an optimal behavior that accounts for model un- certainty in dynamical systems. However, pre- vious studies have shown that by considering the worst case scenario, robust policies can be overly conservative. Our soft-robust framework is an attempt to overcome this issue. In this paper, we present a novel Soft-Robust Actor- Critic algorithm (SR-AC). It learns an optimal policy with respect to a distribution over an uncertainty set and stays robust to model uncer- tainty but avoids the conservativeness of robust strategies. We show the convergence of SR-AC and test the efficiency of our approach on dif- ferent domains by comparing it against regular learning methods and their robust formulations.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-derman18a, title = {Soft-Robust Actor-Critic Policy-Gradient}, author = {Derman, Esther and Mankowitz, Daniel J and Mann, Timothy A and Mannor, Shie}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {207--217}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/derman18a/derman18a.pdf}, url = {https://proceedings.mlr.press/r16/derman18a.html}, abstract = {Robust Reinforcement Learning aims to derive an optimal behavior that accounts for model un- certainty in dynamical systems. However, pre- vious studies have shown that by considering the worst case scenario, robust policies can be overly conservative. Our soft-robust framework is an attempt to overcome this issue. In this paper, we present a novel Soft-Robust Actor- Critic algorithm (SR-AC). It learns an optimal policy with respect to a distribution over an uncertainty set and stays robust to model uncer- tainty but avoids the conservativeness of robust strategies. We show the convergence of SR-AC and test the efficiency of our approach on dif- ferent domains by comparing it against regular learning methods and their robust formulations.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Soft-Robust Actor-Critic Policy-Gradient %A Esther Derman %A Daniel J Mankowitz %A Timothy A Mann %A Shie Mannor %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-derman18a %I PMLR %P 207--217 %U https://proceedings.mlr.press/r16/derman18a.html %V R16 %X Robust Reinforcement Learning aims to derive an optimal behavior that accounts for model un- certainty in dynamical systems. However, pre- vious studies have shown that by considering the worst case scenario, robust policies can be overly conservative. Our soft-robust framework is an attempt to overcome this issue. In this paper, we present a novel Soft-Robust Actor- Critic algorithm (SR-AC). It learns an optimal policy with respect to a distribution over an uncertainty set and stays robust to model uncer- tainty but avoids the conservativeness of robust strategies. We show the convergence of SR-AC and test the efficiency of our approach on dif- ferent domains by comparing it against regular learning methods and their robust formulations. %Z Reissued by PMLR on 04 October 2026.
APA
Derman, E., Mankowitz, D.J., Mann, T.A. & Mannor, S.. (2018). Soft-Robust Actor-Critic Policy-Gradient. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:207-217 Available from https://proceedings.mlr.press/r16/derman18a.html. Reissued by PMLR on 04 October 2026.

Related Material