Influences of Neural Network Structures on an Efficient Reinforcement Learning Policy Search
Wangshu Zhu, André Rosendo · 2019
The Black-box Data-efficient RObot Policy Search algorithm, also known as Black-DROPS, is one of the most data-efficient algorithms for Reinforcement Learning for robotics. The algorithm is based on studying the dynamical model of uncertainty of robots, learning and optimizing the corresponding policy to maximize the reward. Black-DROPS does not limit the reward function, so it can have a wide range of applications. The algorithm is an optimization algorithm using the numerical approach, which is more efficient than the analytical approach in the case of multi-core operation. But the default one layer neural network does not have enough power to meet some difficult problems, and when faced with a high dimensional problem the algorithm will take an unusually long time to solve it. In this paper, we explore different Neural Network structures (layers, neurons, and computer cores) to study their influence on the speed from Black-DROPS on a Double Inverted Pendulum simulation. We demonstrate that although the performance (number of iterations to reach optimality) is affected by structural changes to a lesser extent, the computation time (time per iteration) is drastically affected by such changes, and an optimal structure size seems to exist, as we highlight in our results.