BSO-ES: A Hybrid Direct Policy Search Algorithm for Reinforcement Learning
Youkui Zhang, Liang Zhang, Qiqi Duan, Yuhui Shi · 2021 IEEE Symposium Series on Computational Intelligence (SSCI) · 2021
Recently, it has been demonstrated that evolution strategies (ES), a class of blackbox optimization algorithms, can achieve competitive performance on many reinformance learning (RL) tasks, compared to deep RL algorithms. However, ES can be seen as a randomized local search algorithm based on stochastic finite difference and it is often hard to escape from deep or deceptive local optimum, which means that its exploration ability can be further improved. Brain storm optimization (BSO) is a swarm intelligence algorithm with tradeoff between exploration and exploitation via clustering. To enjoy best of both worlds, this paper proposes a hybrid direct policy search algorithm (BSO-ES) for RL tasks within a distributed computing framework. Specifically, we use BSO as a global searcher to explore the parameter space and ES as a local searcher to further fine-tune policies. We maintain multiple parallel policies at the same time and exchange preferable policies periodically. Experimental results on six challenging continuous control tasks from MuJoCo show that our proposed algorithm is superior to or competitive with OpenAI-ES. In particular, our algorithm has less failures to reach a fixed reward threshold.