Evolutionary reinforcement learning for sparse rewards

Shibei Zhu, Francesco Belardinelli, Borja G. León · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 2021

Temporal logic (TL) is an expressive way of specifying complex goals in reinforcement learning (RL), which facilitates the design of reward functions. However, the combination of these two techniques is prone to generate sparse rewards, which might hinder the learning process. Evolutionary algorithms (EAs) hold promise in tackling this problem by encouraging the diversification of policies through exploration in the parameter space. In this paper, we present GEATL, the first hybrid on-policy evolutionary-based algorithm that combines the advantages of gradient learning in deep RL with the exploration ability of evolutionary algorithms, in order to solve the sparse reward problem pertaining to TL specifications. We test our approach in a delayed reward scenario. Differently from previous baselines combining RL and TL, we show that GEATL is able to tackle complex TL specifications even in sparse-reward settings.

Read the paper · More papers on PaperTik