A Scalable Method for Solving High-Dimensional Continuous POMDPs Using Local Approximation

Tom Erez, William Smart
Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, PMLR R8:176-183, 2010.

Abstract

Partially-Observable Markov Decision Processes (POMDPs) are typically solved by finding an approximate global solution to a corresponding belief-MDP. In this paper, we offer a new plan- ning algorithm for POMDPs with continuous state, action and observation spaces. Since such domains have an inherent notion of locality, we can find an approximate solution using local op- timization methods. We parameterize the belief distribution as a Gaussian mixture, and use the Extended Kalman Filter (EKF) to approximate the belief update. Since the EKF is a first-order filter, we can marginalize over the observations analytically. By using feedback control and state estimation during policy execution, we recover a behavior that is effectively conditioned on in- coming observations despite the unconditioned planning. Local optimization provides no guar- antees of global optimality, but it allows us to tackle domains that are at least an order of mag- nitude larger than the current state-of-the-art. We demonstrate the scalability of our algorithm by considering a simulated hand-eye coordination domain with 16 continuous state dimensions and 6 continuous action dimensions.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR8-erez10a, title = {A Scalable Method for Solving High-Dimensional Continuous POMDPs Using Local Approximation}, author = {Erez, Tom and Smart, William}, booktitle = {Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence}, pages = {176--183}, year = {2010}, editor = {Grünwald, Peter and Spirtes, Peter}, volume = {R8}, series = {Proceedings of Machine Learning Research}, month = {08--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r8/main/assets/erez10a/erez10a.pdf}, url = {https://proceedings.mlr.press/r8/erez10a.html}, abstract = {Partially-Observable Markov Decision Processes (POMDPs) are typically solved by finding an approximate global solution to a corresponding belief-MDP. In this paper, we offer a new plan- ning algorithm for POMDPs with continuous state, action and observation spaces. Since such domains have an inherent notion of locality, we can find an approximate solution using local op- timization methods. We parameterize the belief distribution as a Gaussian mixture, and use the Extended Kalman Filter (EKF) to approximate the belief update. Since the EKF is a first-order filter, we can marginalize over the observations analytically. By using feedback control and state estimation during policy execution, we recover a behavior that is effectively conditioned on in- coming observations despite the unconditioned planning. Local optimization provides no guar- antees of global optimality, but it allows us to tackle domains that are at least an order of mag- nitude larger than the current state-of-the-art. We demonstrate the scalability of our algorithm by considering a simulated hand-eye coordination domain with 16 continuous state dimensions and 6 continuous action dimensions.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T A Scalable Method for Solving High-Dimensional Continuous POMDPs Using Local Approximation %A Tom Erez %A William Smart %B Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2010 %E Peter Grünwald %E Peter Spirtes %F pmlr-vR8-erez10a %I PMLR %P 176--183 %U https://proceedings.mlr.press/r8/erez10a.html %V R8 %X Partially-Observable Markov Decision Processes (POMDPs) are typically solved by finding an approximate global solution to a corresponding belief-MDP. In this paper, we offer a new plan- ning algorithm for POMDPs with continuous state, action and observation spaces. Since such domains have an inherent notion of locality, we can find an approximate solution using local op- timization methods. We parameterize the belief distribution as a Gaussian mixture, and use the Extended Kalman Filter (EKF) to approximate the belief update. Since the EKF is a first-order filter, we can marginalize over the observations analytically. By using feedback control and state estimation during policy execution, we recover a behavior that is effectively conditioned on in- coming observations despite the unconditioned planning. Local optimization provides no guar- antees of global optimality, but it allows us to tackle domains that are at least an order of mag- nitude larger than the current state-of-the-art. We demonstrate the scalability of our algorithm by considering a simulated hand-eye coordination domain with 16 continuous state dimensions and 6 continuous action dimensions. %Z Reissued by PMLR on 04 October 2026.
APA
Erez, T. & Smart, W.. (2010). A Scalable Method for Solving High-Dimensional Continuous POMDPs Using Local Approximation. Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R8:176-183 Available from https://proceedings.mlr.press/r8/erez10a.html. Reissued by PMLR on 04 October 2026.

Related Material