Evolving Network Topology in Policy Gradient Reinforcement Learning Algorithms

Abhijit Banerjee, Dipendranath Ghosh, Suvrojit Das · 2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP) · 2019

In this study we examine the influence of incorporating Evolutionary Algorithms (JADE) and Metaheuristic Algorithms (PSO) for evolving neural network topologies in simple policy gradient, A2C and Deep Deterministic Policy Gradient (DDPG) reinforcement learning. A2C is experimented on a discrete action space model free environment followed by DDPG on a continious action model free environment;“LunarlanderContinious-v2” from OpenAI gym. This baseline result is then compared with topology and hyperparameter optimized A2C and DDPG using Differential Evolution and Particle Swarm Optimizaton. The results thus obtained shows that Differential Evolution (JADE) optimized DDPG has the fastest convergence and minimum forgetting of learned weights.

Read the paper · More papers on PaperTik