Understanding the Role of Population Experiences in Proximal Distilled Evolutionary Reinforcement Learning

Thai Huy Nguyen, Ngoc Hoang Luong · 2023

Evolutionary Reinforcement Learning (ERL) combines the sample-efficiency property of Reinforcement Learning and exploration capabilities from the population-based search of Evolutionary Computation. These methods have shown promising performance on many continuous control tasks. However, one could observe the instability that may occur from such methods. Several works have shown that the experiences coming from the population individuals lead the state distribution shift in the RL policy updating process. A vanilla remedy method has been proposed to alleviate this issue by separating the experience transitions into two distinct replay buffers for the RL policy and the population and mixing the samples from the two buffers with a fixed ratio to update the RL policy. The effectiveness of this approach has been shown empirically on an ERL method where Evolution Strategies (ES) assists an external RL agent. Nevertheless, there has not been any thorough investigation on Genetic Algorithm (GA) based ERL to understand how this method performs on these ERL approaches. In this paper, we analyze the influence of off-policy data coming from the GA population to the RL policy and how the mixing method performs on a state-of-the-art ERL method, namely Proximal Distilled Evolutionary Reinforcement Learning (PDERL).

Read the paper · More papers on PaperTik