On the effect of the sampling ratio of past trajectories in the combination of evolutionary algorithm and deep reinforcement learning

Yasen Wang, Youhei Akimoto · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 2022

Combinations of evolutionary approaches (EA) and deep reinforcement learning approaches (DRL) have been proposed to take advantages of two approaches: stability of EA and sample efficiency of DRL. CEM-RL is a promising approach combining cross-entropy method (CEM) and twin delayed deep deterministic policy gradient (TD3), showing competitive performance in this fields. In the TD3 part, interaction histories obtained by CEM-oriented agents and TD3-oriented agents are mixed in a single replay buffer, and mini-batch samples are drawn uniform-randomly from the replay buffer. The quality of the training in the TD3 part must depend on the distribution of samples in the replay buffer, hence it depends on the ratio of CEM-oriented and TD3-oriented samples, called replay sample ratio. In this paper we investigate the impact of the replay sample ratio in CEM-RL on 5 MuJoCo environments and confirm that a good ratio leads to outperforming the original CEM-RL on some environments. Based on this observation, we propose an adaptation mechanism for the replay sample ratio, and empirically show that the proposed variant, RSR-CEM-RL, outperforms the original CEM-RL on some environments.

Read the paper · More papers on PaperTik