Exploration and Exploitation in Reinforcement Learning
Felix Adelsbo · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2018
In reinforcement learning there exists a dilemma of exploration versus exploitation. This has led to the development of methods that have di↵er- ent approaches to this. Using di↵erent types of methods, and modifying them in di↵erent ways can lead to di↵erent results. Knowledge of how di↵erent methods work can give knowledge of what should be used in aspecific case. Two ways that methods can be modified are change in adjustable pa- rameters and change in the number of steps of random actions at thebeginning. How much these two modifications e↵ect the results in a spe- cific environment may di↵er a lot, and can be a very critical thing to consider for certain results. The goal of this study is to answer the question of how the performance of the di↵erent methods such as random, greedy, ✏-greedy, ✏-decreasingand Softmax is a↵ected by di↵erent values of their adjustable parameters,and by the number of steps of random actions at the beginning. Thesimulations and a comparative analysis are conducted for the case of an inverted pendulum with a vertical pole placed on a moving cart.