Approximate Kalman Filter Q-Learning for Continuous State-Space MDPs

Charles Tripp, Ross Shachter
Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence, PMLR R11:687-696, 2013.

Abstract

We seek to learn an effective policy for a Markov Decision Process (MDP) with con- tinuous states via Q-Learning. Given a set of basis functions over state action pairs we search for a corresponding set of linear weights that minimizes the mean Bellman residual. Our algorithm uses a Kalman filter model to estimate those weights and we have developed a simpler approximate Kalman fil- ter model that outperforms the current state of the art projected TD-Learning methods on several standard benchmark problems.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR11-tripp13a, title = {Approximate {K}alman Filter Q-Learning for Continuous State-Space MDPs}, author = {Tripp, Charles and Shachter, Ross}, booktitle = {Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence}, pages = {687--696}, year = {2013}, editor = {Nicholson, Ann and Smyth, Padhraic}, volume = {R11}, series = {Proceedings of Machine Learning Research}, month = {12--14 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r11/main/assets/tripp13a/tripp13a.pdf}, url = {https://proceedings.mlr.press/r11/tripp13a.html}, abstract = {We seek to learn an effective policy for a Markov Decision Process (MDP) with con- tinuous states via Q-Learning. Given a set of basis functions over state action pairs we search for a corresponding set of linear weights that minimizes the mean Bellman residual. Our algorithm uses a Kalman filter model to estimate those weights and we have developed a simpler approximate Kalman fil- ter model that outperforms the current state of the art projected TD-Learning methods on several standard benchmark problems.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Approximate Kalman Filter Q-Learning for Continuous State-Space MDPs %A Charles Tripp %A Ross Shachter %B Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2013 %E Ann Nicholson %E Padhraic Smyth %F pmlr-vR11-tripp13a %I PMLR %P 687--696 %U https://proceedings.mlr.press/r11/tripp13a.html %V R11 %X We seek to learn an effective policy for a Markov Decision Process (MDP) with con- tinuous states via Q-Learning. Given a set of basis functions over state action pairs we search for a corresponding set of linear weights that minimizes the mean Bellman residual. Our algorithm uses a Kalman filter model to estimate those weights and we have developed a simpler approximate Kalman fil- ter model that outperforms the current state of the art projected TD-Learning methods on several standard benchmark problems. %Z Reissued by PMLR on 04 October 2026.
APA
Tripp, C. & Shachter, R.. (2013). Approximate Kalman Filter Q-Learning for Continuous State-Space MDPs. Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R11:687-696 Available from https://proceedings.mlr.press/r11/tripp13a.html. Reissued by PMLR on 04 October 2026.

Related Material