Message-passing algorithms for large structured decentralized POMDPs
Akshat Kumar, Shlomo Zilberstein · 2011
Decentralized POMDPs provide a rigorous framework for multi-agent decision-theoretic planning. However, their high complexity has limited scalability. In this work, we present a promising new class of algorithms based on probabilis-tic inference for infinite-horizon ND-POMDPs—a restricted Dec-POMDP model. We first transform the policy opti-mization problem to that of likelihood maximization in a mixture of dynamic Bayes nets (DBNs). We then develop the Expectation-Maximization (EM) algorithm for maximiz-ing the likelihood in this representation. The EM algorithm for ND-POMDPs lends itself naturally to a simple message-passing paradigm guided by the agent interaction graph. It is thus highly scalable w.r.t. the number of agents, can be easily parallelized, and produces good quality solutions.