Efficient Inference in Markov Control Problems

Thomas Furmston, David Barber
Proceedings of the 27th Conference on Uncertainty in Artificial Intelligence, PMLR R9:259-267, 2011.

Abstract

Markov control algorithms that perform smooth, non-greedy updates of the policy have been shown to be very general and versatile, with policy gradient and Expectation Maximisation algorithms being particularly popular. For these algorithms, marginal inference of the reward weighted trajectory distribution is required to perform policy updates. We discuss a new exact inference algorithm for these marginals in the finite horizon case that is more efficient than the standard approach based on classical forward-backward recursions. We also provide a principled extension to infinite horizon Markov Decision Problems that explicitly accounts for an infinite horizon. This extension provides a novel algorithm for both policy gradients and Expectation Maximisation in infinite horizon problems.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR9-furmston11a, title = {Efficient Inference in {M}arkov Control Problems}, author = {Furmston, Thomas and Barber, David}, booktitle = {Proceedings of the 27th Conference on Uncertainty in Artificial Intelligence}, pages = {259--267}, year = {2011}, editor = {Cozman, Fabio and Pfeffer, Avi}, volume = {R9}, series = {Proceedings of Machine Learning Research}, month = {14--17 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r9/main/assets/furmston11a/furmston11a.pdf}, url = {https://proceedings.mlr.press/r9/furmston11a.html}, abstract = {Markov control algorithms that perform smooth, non-greedy updates of the policy have been shown to be very general and versatile, with policy gradient and Expectation Maximisation algorithms being particularly popular. For these algorithms, marginal inference of the reward weighted trajectory distribution is required to perform policy updates. We discuss a new exact inference algorithm for these marginals in the finite horizon case that is more efficient than the standard approach based on classical forward-backward recursions. We also provide a principled extension to infinite horizon Markov Decision Problems that explicitly accounts for an infinite horizon. This extension provides a novel algorithm for both policy gradients and Expectation Maximisation in infinite horizon problems.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Efficient Inference in Markov Control Problems %A Thomas Furmston %A David Barber %B Proceedings of the 27th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2011 %E Fabio Cozman %E Avi Pfeffer %F pmlr-vR9-furmston11a %I PMLR %P 259--267 %U https://proceedings.mlr.press/r9/furmston11a.html %V R9 %X Markov control algorithms that perform smooth, non-greedy updates of the policy have been shown to be very general and versatile, with policy gradient and Expectation Maximisation algorithms being particularly popular. For these algorithms, marginal inference of the reward weighted trajectory distribution is required to perform policy updates. We discuss a new exact inference algorithm for these marginals in the finite horizon case that is more efficient than the standard approach based on classical forward-backward recursions. We also provide a principled extension to infinite horizon Markov Decision Problems that explicitly accounts for an infinite horizon. This extension provides a novel algorithm for both policy gradients and Expectation Maximisation in infinite horizon problems. %Z Reissued by PMLR on 04 October 2026.
APA
Furmston, T. & Barber, D.. (2011). Efficient Inference in Markov Control Problems. Proceedings of the 27th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R9:259-267 Available from https://proceedings.mlr.press/r9/furmston11a.html. Reissued by PMLR on 04 October 2026.

Related Material