[edit]
Anytime Planning for Decentralized POMDPs using Expectation Maximization
Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, PMLR R8:301-308, 2010.
Abstract
Decentralized POMDPs provide an expres- sive framework for multi-agent sequential de- cision making. While finite-horizon DEC- POMDPs have enjoyed significant success, progress remains slow for the infinite-horizon case mainly due to the inherent complexity of optimizing stochastic controllers representing agent policies. We present a promising new class of algorithms for the infinite-horizon case, which recasts the optimization problem as inference in a mixture of DBNs. An attrac- tive feature of this approach is the straight- forward adoption of existing inference tech- niques in DBNs for solving DEC-POMDPs and supporting richer representations such as factored or continuous states and actions. We also derive the Expectation Maximization (EM) algorithm to optimize the joint pol- icy represented as DBNs. Experiments on benchmark domains show that EM compares favorably against the state-of-the-art solvers.