EVOLUTION OF REINFORCEMENT LEARNING AGENTS USING THE GENETIC ALGORITHM
Artem Volokyta, Bohdan Hereha · TECHNICAL SCIENCES AND TECHNOLOGIES · 2023
Reinforcement learning (RL) allows agents to make decisions based on a reward function. However, in the process of learning, the choice of the values of the parameters of the learning algorithm can significantly affect the overall learning process. Agents using the policy gradient algorithm can be trained for a long time, but even then, they may not behave perfectly Thinking more about it, we realized that the reason for the long training is that gradients are almost absent, and therefore not very useful. Gradients help in supervised learning tasks, such as image classification, by providing useful information on howto change the parameters (weights or offsets) of the network for better accuracy. In image classification, after each mini-series of training, backpropagation provides a clear gradient (direction) for each parameter in the network. In reinforcement learning, however, the gradient information is only provided occasionally when the environment provides a reward or punishment.In most cases, our agent performs actions without knowing whether they are useful or not. Therefore, in this paper, we will improve the agents by using a genetic algorithm, i.e., we evolve the agents.This research explores the use of genetic algorithms to improve the performance of reinforcement learning agents. We conducted a series of trials using various neural network parameters, including weights, biases, and activation functions, in order to find the optimal values that cause the agent to receive more rewards. Our approach includes the use of domain knowledge to initialize the population of the genetic algorithm as well as to evaluate solutions. This allows us to direct the search towards more promising solutions. Special attention is paid to the impact of various genetic algorithm parameters onlearning efficiency. The potential applications of this research are broad, ranging from robotics and autonomous vehicles to gaming and finance. The results of the study can also be used to develop new algorithms and methods to improve the performance of reinforcement learning agents, which further contributes to the development of machine learning.Our research has shown that the use of a genetic algorithm can significantly improve the efficiency of agent learning.The result is the successful completion of the CartPole-v0 game by evolved agents. 98 % of our population will reach the maximum, i.e. successfully complete the game.