Reinforcement Learning and the Parallel Actions Common Goal problem

Bachelor Thesis, Ivan Koster, Martijn van Otterlo · 2012

There exists an interesting category of problems, in particular in real-time strategy and colonization games. The player can perform several actions, usually in parallel, like gathering resources and building structures, with a single goal in mind. He has to consider which actions to take and how to schedule them, in order to reach the goal in the fastest way possible. We call these problems parallel actions common goal (PACG) problems. In this thesis we will try to solve these problems with reinforcement learning. To do this, we show that it is possible to capture PACG problems in a Markov Decision Processes (MDP). An MDP can be translated into a programming language, resulting in a simulator. Next, the simulator can be used as the environment for a reinforcement learning agent. We will focus on one complex PACG problem in particular, the browser game Ogame. After that, we will show that, if implemented in a tabular form, Q-learning can consistently learn the optimal policy in simple PACG problems. This will be done using a simplified version of the Ogame problem. Following this, we will introduce a Qlearning algorithm for more complex PACG problems, which includes Artificial Neural Networks and Experience Replay. We will then test this algorithm on the simplified Ogame problem to show that it also learns optimal policies. At last, we will try to use this algorithm to learn good policies for the complex Ogame problem.

Read the paper · More papers on PaperTik