New inference strategies for solving Markov Decision Processes using reversible jump MCMC

Matt Hoffman, Hendrik Kueck, Nando de Freitas, Arnaud Doucet
Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, PMLR R7:223-231, 2009.

Abstract

In this paper we build on previous work which uses inferences techniques, in particular Markov Chain Monte Carlo (MCMC) methods, to solve parameterized control problems. We propose a number of modifications in order to make this approach more practical in general, higher-dimensional spaces. We first introduce a new target distribution which is able to incorporate more reward information from sampled trajectories. We also show how to break strong correlations between the policy parameters and sampled trajectories in order to sample more freely. Finally, we show how to incorporate these techniques in a principled manner to obtain estimates of the optimal policy.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR7-hoffman09a, title = {New inference strategies for solving {M}arkov Decision Processes using reversible jump {MCMC}}, author = {Hoffman, Matt and Kueck, Hendrik and de Freitas, Nando and Doucet, Arnaud}, booktitle = {Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence}, pages = {223--231}, year = {2009}, editor = {Bilmes, Jeff and Ng, Andrew Y.}, volume = {R7}, series = {Proceedings of Machine Learning Research}, month = {18--21 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r7/main/assets/hoffman09a/hoffman09a.pdf}, url = {https://proceedings.mlr.press/r7/hoffman09a.html}, abstract = {In this paper we build on previous work which uses inferences techniques, in particular Markov Chain Monte Carlo (MCMC) methods, to solve parameterized control problems. We propose a number of modifications in order to make this approach more practical in general, higher-dimensional spaces. We first introduce a new target distribution which is able to incorporate more reward information from sampled trajectories. We also show how to break strong correlations between the policy parameters and sampled trajectories in order to sample more freely. Finally, we show how to incorporate these techniques in a principled manner to obtain estimates of the optimal policy.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T New inference strategies for solving Markov Decision Processes using reversible jump MCMC %A Matt Hoffman %A Hendrik Kueck %A Nando de Freitas %A Arnaud Doucet %B Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2009 %E Jeff Bilmes %E Andrew Y. Ng %F pmlr-vR7-hoffman09a %I PMLR %P 223--231 %U https://proceedings.mlr.press/r7/hoffman09a.html %V R7 %X In this paper we build on previous work which uses inferences techniques, in particular Markov Chain Monte Carlo (MCMC) methods, to solve parameterized control problems. We propose a number of modifications in order to make this approach more practical in general, higher-dimensional spaces. We first introduce a new target distribution which is able to incorporate more reward information from sampled trajectories. We also show how to break strong correlations between the policy parameters and sampled trajectories in order to sample more freely. Finally, we show how to incorporate these techniques in a principled manner to obtain estimates of the optimal policy. %Z Reissued by PMLR on 04 October 2026.
APA
Hoffman, M., Kueck, H., de Freitas, N. & Doucet, A.. (2009). New inference strategies for solving Markov Decision Processes using reversible jump MCMC. Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R7:223-231 Available from https://proceedings.mlr.press/r7/hoffman09a.html. Reissued by PMLR on 04 October 2026.

Related Material