Experimental Demonstration of Adaptive MDP-Based Planning with Model Uncertainty
Brett Bethke, Luca F. Bertuccelli, Jonathan P. How · AIAA Guidance, Navigation, and Control Conference and Exhibit · 2008
Markov decision processes (MDPs) are a natural framework for solving multiagent planning problems since they can model stochastic system dynamics and interdependencies between agents. In these approaches, accurate modeling of the system in question is important, since mismodeling may lead to severely degraded performance (i.e. loss of vehicles). Furthermore, in many problems of interest, it may be difficult or impossible to obtain an accurate model before the system begins operating; rather, the model must be estimated online. Therefore, an adaptation mechanism that can estimate the system model and adjust the system control policy online can improve performance over a static (off-line) approach. This paper presents an MDP formulation of a multi-agent persistent surveillance problem and shows, in simulation, the importance of accurate modeling of the system. An adaptation mechanism, consisting of a Bayesian model estimator and a continuouslyrunning MDP solver, is then discussed. Finally, we present hardware flight results from the MIT RAVEN testbed that clearly demonstrate the performance benefits of this adaptive approach in the persistent surveillance problem. I.