Stein Variational Policy Gradient

Yang Liu, Prajit Ramachandran, Qiang Liu, Jian Peng
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:721-729, 2017.

Abstract

Policy gradient methods have been successfully applied to many complex reinforcement learn- ing problems. However, policy gradient meth- ods suffer from high variance, slow convergence, and inefficient exploration. In this work, we in- troduce a maximum entropy policy optimization framework which explicitly encourages param- eter exploration, and show that this framework can be reduced to a Bayesian inference problem. We then propose a novel Stein variational pol- icy gradient method (SVPG) which combines ex- isting policy gradient methods and a repulsive functional to generate a set of diverse but well- behaved policies. SVPG is robust to random ini- tializations and can easily be implemented in a parallel manner. On several continuous control problems, we find that SVPG versions of RE- INFORCE and advantage actor-critic algorithms are greatly improved in terms of both average re- turn and data efficiency.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR15-liu17b, title = {{S}tein Variational Policy Gradient}, author = {Liu, Yang and Ramachandran, Prajit and Liu, Qiang and Peng, Jian}, booktitle = {Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence}, pages = {721--729}, year = {2017}, editor = {Elidan, Gal and Kersting, Kristian}, volume = {R15}, series = {Proceedings of Machine Learning Research}, month = {11--15 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r15/main/assets/liu17b/liu17b.pdf}, url = {https://proceedings.mlr.press/r15/liu17b.html}, abstract = {Policy gradient methods have been successfully applied to many complex reinforcement learn- ing problems. However, policy gradient meth- ods suffer from high variance, slow convergence, and inefficient exploration. In this work, we in- troduce a maximum entropy policy optimization framework which explicitly encourages param- eter exploration, and show that this framework can be reduced to a Bayesian inference problem. We then propose a novel Stein variational pol- icy gradient method (SVPG) which combines ex- isting policy gradient methods and a repulsive functional to generate a set of diverse but well- behaved policies. SVPG is robust to random ini- tializations and can easily be implemented in a parallel manner. On several continuous control problems, we find that SVPG versions of RE- INFORCE and advantage actor-critic algorithms are greatly improved in terms of both average re- turn and data efficiency.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Stein Variational Policy Gradient %A Yang Liu %A Prajit Ramachandran %A Qiang Liu %A Jian Peng %B Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2017 %E Gal Elidan %E Kristian Kersting %F pmlr-vR15-liu17b %I PMLR %P 721--729 %U https://proceedings.mlr.press/r15/liu17b.html %V R15 %X Policy gradient methods have been successfully applied to many complex reinforcement learn- ing problems. However, policy gradient meth- ods suffer from high variance, slow convergence, and inefficient exploration. In this work, we in- troduce a maximum entropy policy optimization framework which explicitly encourages param- eter exploration, and show that this framework can be reduced to a Bayesian inference problem. We then propose a novel Stein variational pol- icy gradient method (SVPG) which combines ex- isting policy gradient methods and a repulsive functional to generate a set of diverse but well- behaved policies. SVPG is robust to random ini- tializations and can easily be implemented in a parallel manner. On several continuous control problems, we find that SVPG versions of RE- INFORCE and advantage actor-critic algorithms are greatly improved in terms of both average re- turn and data efficiency. %Z Reissued by PMLR on 04 October 2026.
APA
Liu, Y., Ramachandran, P., Liu, Q. & Peng, J.. (2017). Stein Variational Policy Gradient. Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R15:721-729 Available from https://proceedings.mlr.press/r15/liu17b.html. Reissued by PMLR on 04 October 2026.

Related Material