Rollout Sampling Policy Iteration for Decentralized POMDPs

Feng Wu, Shlomo Zilberstein, Xiaoping Chen
Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, PMLR R8:665-672, 2010.

Abstract

We present decentralized rollout sampling pol- icy iteration (DecRSPI) — a new algorithm for multi-agent decision problems formalized as DEC-POMDPs. DecRSPI is designed to im- prove scalability and tackle problems that lack an explicit model. The algorithm uses Monte- Carlo methods to generate a sample of reachable belief states. Then it computes a joint policy for each belief state based on the rollout estimations. A new policy representation allows us to repre- sent solutions compactly. The key benefits of the algorithm are its linear time complexity over the number of agents, its bounded memory usage and good solution quality. It can solve larger prob- lems that are intractable for existing planning al- gorithms. Experimental results confirm the ef- fectiveness and scalability of the approach.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR8-wu10a, title = {Rollout Sampling Policy Iteration for Decentralized POMDPs}, author = {Wu, Feng and Zilberstein, Shlomo and Chen, Xiaoping}, booktitle = {Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence}, pages = {665--672}, year = {2010}, editor = {Grünwald, Peter and Spirtes, Peter}, volume = {R8}, series = {Proceedings of Machine Learning Research}, month = {08--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r8/main/assets/wu10a/wu10a.pdf}, url = {https://proceedings.mlr.press/r8/wu10a.html}, abstract = {We present decentralized rollout sampling pol- icy iteration (DecRSPI) — a new algorithm for multi-agent decision problems formalized as DEC-POMDPs. DecRSPI is designed to im- prove scalability and tackle problems that lack an explicit model. The algorithm uses Monte- Carlo methods to generate a sample of reachable belief states. Then it computes a joint policy for each belief state based on the rollout estimations. A new policy representation allows us to repre- sent solutions compactly. The key benefits of the algorithm are its linear time complexity over the number of agents, its bounded memory usage and good solution quality. It can solve larger prob- lems that are intractable for existing planning al- gorithms. Experimental results confirm the ef- fectiveness and scalability of the approach.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Rollout Sampling Policy Iteration for Decentralized POMDPs %A Feng Wu %A Shlomo Zilberstein %A Xiaoping Chen %B Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2010 %E Peter Grünwald %E Peter Spirtes %F pmlr-vR8-wu10a %I PMLR %P 665--672 %U https://proceedings.mlr.press/r8/wu10a.html %V R8 %X We present decentralized rollout sampling pol- icy iteration (DecRSPI) — a new algorithm for multi-agent decision problems formalized as DEC-POMDPs. DecRSPI is designed to im- prove scalability and tackle problems that lack an explicit model. The algorithm uses Monte- Carlo methods to generate a sample of reachable belief states. Then it computes a joint policy for each belief state based on the rollout estimations. A new policy representation allows us to repre- sent solutions compactly. The key benefits of the algorithm are its linear time complexity over the number of agents, its bounded memory usage and good solution quality. It can solve larger prob- lems that are intractable for existing planning al- gorithms. Experimental results confirm the ef- fectiveness and scalability of the approach. %Z Reissued by PMLR on 04 October 2026.
APA
Wu, F., Zilberstein, S. & Chen, X.. (2010). Rollout Sampling Policy Iteration for Decentralized POMDPs. Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R8:665-672 Available from https://proceedings.mlr.press/r8/wu10a.html. Reissued by PMLR on 04 October 2026.

Related Material