Anytime Planning for Decentralized POMDPs using Expectation Maximization

Akshat Kumar, Shlomo Zilberstein
Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, PMLR R8:301-308, 2010.

Abstract

Decentralized POMDPs provide an expres- sive framework for multi-agent sequential de- cision making. While finite-horizon DEC- POMDPs have enjoyed significant success, progress remains slow for the infinite-horizon case mainly due to the inherent complexity of optimizing stochastic controllers representing agent policies. We present a promising new class of algorithms for the infinite-horizon case, which recasts the optimization problem as inference in a mixture of DBNs. An attrac- tive feature of this approach is the straight- forward adoption of existing inference tech- niques in DBNs for solving DEC-POMDPs and supporting richer representations such as factored or continuous states and actions. We also derive the Expectation Maximization (EM) algorithm to optimize the joint pol- icy represented as DBNs. Experiments on benchmark domains show that EM compares favorably against the state-of-the-art solvers.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR8-kumar10a, title = {Anytime Planning for Decentralized POMDPs using Expectation Maximization}, author = {Kumar, Akshat and Zilberstein, Shlomo}, booktitle = {Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence}, pages = {301--308}, year = {2010}, editor = {Grünwald, Peter and Spirtes, Peter}, volume = {R8}, series = {Proceedings of Machine Learning Research}, month = {08--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r8/main/assets/kumar10a/kumar10a.pdf}, url = {https://proceedings.mlr.press/r8/kumar10a.html}, abstract = {Decentralized POMDPs provide an expres- sive framework for multi-agent sequential de- cision making. While finite-horizon DEC- POMDPs have enjoyed significant success, progress remains slow for the infinite-horizon case mainly due to the inherent complexity of optimizing stochastic controllers representing agent policies. We present a promising new class of algorithms for the infinite-horizon case, which recasts the optimization problem as inference in a mixture of DBNs. An attrac- tive feature of this approach is the straight- forward adoption of existing inference tech- niques in DBNs for solving DEC-POMDPs and supporting richer representations such as factored or continuous states and actions. We also derive the Expectation Maximization (EM) algorithm to optimize the joint pol- icy represented as DBNs. Experiments on benchmark domains show that EM compares favorably against the state-of-the-art solvers.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Anytime Planning for Decentralized POMDPs using Expectation Maximization %A Akshat Kumar %A Shlomo Zilberstein %B Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2010 %E Peter Grünwald %E Peter Spirtes %F pmlr-vR8-kumar10a %I PMLR %P 301--308 %U https://proceedings.mlr.press/r8/kumar10a.html %V R8 %X Decentralized POMDPs provide an expres- sive framework for multi-agent sequential de- cision making. While finite-horizon DEC- POMDPs have enjoyed significant success, progress remains slow for the infinite-horizon case mainly due to the inherent complexity of optimizing stochastic controllers representing agent policies. We present a promising new class of algorithms for the infinite-horizon case, which recasts the optimization problem as inference in a mixture of DBNs. An attrac- tive feature of this approach is the straight- forward adoption of existing inference tech- niques in DBNs for solving DEC-POMDPs and supporting richer representations such as factored or continuous states and actions. We also derive the Expectation Maximization (EM) algorithm to optimize the joint pol- icy represented as DBNs. Experiments on benchmark domains show that EM compares favorably against the state-of-the-art solvers. %Z Reissued by PMLR on 04 October 2026.
APA
Kumar, A. & Zilberstein, S.. (2010). Anytime Planning for Decentralized POMDPs using Expectation Maximization. Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R8:301-308 Available from https://proceedings.mlr.press/r8/kumar10a.html. Reissued by PMLR on 04 October 2026.

Related Material