[edit]
Stein Variational Policy Gradient
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:721-729, 2017.
Abstract
Policy gradient methods have been successfully applied to many complex reinforcement learn- ing problems. However, policy gradient meth- ods suffer from high variance, slow convergence, and inefficient exploration. In this work, we in- troduce a maximum entropy policy optimization framework which explicitly encourages param- eter exploration, and show that this framework can be reduced to a Bayesian inference problem. We then propose a novel Stein variational pol- icy gradient method (SVPG) which combines ex- isting policy gradient methods and a repulsive functional to generate a set of diverse but well- behaved policies. SVPG is robust to random ini- tializations and can easily be implemented in a parallel manner. On several continuous control problems, we find that SVPG versions of RE- INFORCE and advantage actor-critic algorithms are greatly improved in terms of both average re- turn and data efficiency.