Auto-Tuning Online POMDPs for Multi-Object Search in Uncertain Environments

Shohei Wakayama, Nisar R. Ahmed · AIAA Scitech 2020 Forum · 2020

Future exploration missions will require mobile robots to autonomously solve the challenging problem of searching for and localizing multiple objects of interest in uncertain environments. Such problems can be posed as partially observable Markov decision processes (POMDPs), which require approximations to obtain practically implementable policies for sequential decision making under uncertainty. Although online solvers such as Partially Observable Monte Carlo Planning (POMCP) have recently attracted much interest due to their fast anytime closed-loop capabilities, they require extensive hyperparameter tuning before implementation to achieve acceptable performance. However, in missions where robots cannot receive any external aid for long periods of time, they must be able to adjust and adapt hyperparameters to varying operating conditions and environmental contexts. In this paper, we cast this hyperparameter selection problem for search-tree based online solvers a Multi-Armed Bandit (MAB). Our contributions for this problem are as follows. First of all, we evaluate which existing MAB policy approximation strategies are suitable for this problem and examine the definition of utility function for MAB auto-tuning. Then, the performance of auto-tuning robot is evaluated in a simple baseline fully observable Markov Decision Process (MDP) single goal navigation problem using Monte Carlo Tree Search policy approximation. It is shown that the performance of the robot using auto-tuned MCTS planner is shown to be almost the same as the best theoretical performance produced by offline value iteration. Next, the auto-tuning robot is evaluated in a simulated fully observable multi-goal navigation problem in more complex environments, showing that the auto-tuning robot performs equally/better than an offline tuned fixed hyperparameter robot. Finally, the auto-tuning robot is evaluated in simulation for a multi-object search problem in uncertain environment where locations and number of objects are unknown. The results validate that the proposed tree-search auto-tuning process works even in partially observable environments.

Read the paper · More papers on PaperTik