Deep Apprenticeship Learning for Playing Video Games

Miroslav Bogdanović, Dejan Markovikj, Misha Denil, Nando de Freitas · 2015

Recently it has been shown that deep neural networks can learn to play Atari games by directly observing raw pixels of the playing area. We show how apprenticeship learning can be applied in this setting so that an agent can learn to perform a task (i.e. play a game) by observing the expert, without any explicitly provided knowledge of the game’s internal state or objectives. Background Mnih et al. (2013) recently demonstrated that it is possible to combine Q-learning with deep learning to play Atari games. Their method learns to maximize the score of the game, which is explicitly provided to the model during training. We extend the approach of Mnih et al. (2013) to the ap-prenticeship learning setting, allowing our agent to learn to play without being provided with any explicit knowledge of the game score. In this paper we take a very simple approach to apprenticeship learning by simply observing the expert play and training a classifier to identify expert actions from game states. Deep Apprenticeship Learning Following Mnih et al. (2013), we train a convolutional neu-ral network to play Atari games by observing only raw pixels of the playing area. However, instead of learning to maxi-mize the game score directly, we attempt to imitate the be-haviour of an expert player. By watching an expert play, our network is able to learn to map game states to actions in a way that does not require that the score of the game is pro-vided externally. We call our method Deep Apprenticeship Learning (DAL). Data Collection To interact with Atari games we used the Arcade Learn-ing Environment (Naddaf 2010). To collect training data we modified the Arcade Learning Environment to record states (video frames) and actions (button presses) while a human plays the game. We collected human gameplay data for the Freeway game. Freeway is a classical game about trying to cross a street Copyright c © 2015, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved. actions (a) fully connected MLP convolutional layers state

Read the paper · More papers on PaperTik