[edit]
Soft-Robust Actor-Critic Policy-Gradient
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:207-217, 2018.
Abstract
Robust Reinforcement Learning aims to derive an optimal behavior that accounts for model un- certainty in dynamical systems. However, pre- vious studies have shown that by considering the worst case scenario, robust policies can be overly conservative. Our soft-robust framework is an attempt to overcome this issue. In this paper, we present a novel Soft-Robust Actor- Critic algorithm (SR-AC). It learns an optimal policy with respect to a distribution over an uncertainty set and stays robust to model uncer- tainty but avoids the conservativeness of robust strategies. We show the convergence of SR-AC and test the efficiency of our approach on dif- ferent domains by comparing it against regular learning methods and their robust formulations.