Exploratory Hybrid Search in Hierarchical Reinforcement Learning

이상엽 · Seoul National University Open Repository (Seoul National University) · 2020

Balancing exploitation and exploration is a great challenge in many optimization problems.Evolutionary algorithms, such as evolutionary strategies and genetic algorithms, are algorithms inspired by biological evolution.They have been used for various optimization problems, such as combinatorial optimization and continuous optimization.However, evolutionary algorithms lack fine-tuning near local optima; in other words, they lack exploitation power.This drawback can be overcome by hybridization.Hybrid genetic algorithms, or memetic algorithms, are successful examples of hybridization.Although the solution space is exponentially vast in some optimization problems, these algorithms successfully find satisfactory solutions.In the deep learning era, the problem of exploitation and exploration has been relatively neglected.In deep reinforcement learning problems, however, balancing exploitation and exploration is more crucial than that in problems with supervision.Many environments in the real world have an exponentially wide state space that must be explored by agents.Without sufficient exploration power, agents only reveal a small portion of the state space and end up with seeking only instant rewards.In this thesis, a hybridization method is proposed which contains both gradientbased policy optimization with strong exploitation power and evolutionary policy optimization with strong exploration power.First, the gradient-based policy optimization and evolutionary policy optimization are analyzed in various environments.The results demonstrate that evolutionary policy optimization is robust for sparse rewards but weak for instant rewards, whereas gradient-based policy optimization is effective for instant rewards but weak for sparse rewards.This difference between the two optimizations reveals the potential of hybridization in policy optimization.Then, a hybrid search is suggested in the framework of hierarchical reinforcement learning.The results demonstrate that the hybrid search finds an effective agent for complex environments with sparse rewards thanks to its balanced exploitation and exploration.

Read the paper · More papers on PaperTik