Policy Reuse for Learning and Planning in Partially Observable Markov Decision Processes

Bo Wu, Yanpeng Feng · 2017

Learning and planning in partially observable Markov decision processes (POMDPs) is computationally intractable in real-time system. In order to address this problem, this paper proposes a belief policy reuse (BPR) method to avoid repeated computation. Firstly, the policy reuse evaluation mechanism based on belief Kullback-Leibler divergence is presented as a similarity metric between beliefs in the belief-policy library. If the current belief is similar to any of the past ones, the policy in the belief-policy library is reused. Otherwise, BPR exploits Monte-Carlo particle method to explore a new policy, and stores the new policy with belief in the belief-policy library, so it can be reused in the future. The experimental results show that the proposed approach is an effective way for improving the learning efficiency in large-scale partially observable Markov decision processes.

Read the paper · More papers on PaperTik