An evolutionary based approach to profit sharing for POMDP environments

Kohei Suzuki, Shōhei Kato · 2017

In a POMDP environment, an agent may observe the same information at more than one state. HQ-learning and episode-based profit sharing (EPS) are well-known methods for solving this problem. HQ-learning divides a POMDP environment into subtasks. EPS distributes the same reward to state-action pairs in the episode when an agent achieves a goal. However, these methods have disadvantages related to the learning efficiency and localized solutions. In this paper, we propose a hybrid learning method that combines profit sharing and genetic algorithm. We also report the effectiveness of our method by using some experiments with partially observable mazes.

Read the paper · More papers on PaperTik