Learning to play Tetris applying reinforcement learning methods.

Alexander Groß, Jan Friedland, Friedhelm Schwenker · The European Symposium on Artificial Neural Networks · 2008

In this paper the application of reinforcement learning to Tetris is investigated, particulary the idea of temporal difference learning is applied to estimate the state value function V. For two predefined reward functions Tetris agents have been trained by using a � -greedy policy. In the numerical experiments it can be observed that the trained agents can outperform fixed policy agents significantly, e.g. by factor 5 for a complex reward function. Playing games, such like Chess, Go, Checkers, Backgammon or Poker, is always a great intellectual challenge to humans, and therefore game playing is a sce- nario to test and evaluate artificial intelligence methods, in particular machine learning aspects have been taken more and more into account during the last years. Many methods derived from the fields of traditional artificial intelligence and mathematical game theory have been utilized in computer games, for in- stance, game trees are one of the most popular tools. A game tree represents all the possible states of a game in its nodes. Starting point is the root node representing the starting configuration of the game. Children nodes within the tree are representing all states that can be reached for the current node follow- ing the rules of the game, and leaf nodes are representing the terminal states of the game. In almost all (interesting) games the complete tree is too large to search all the possible pathes, and therefore the space to be searched must be reduced by applying heuristics which were tailored by human experts. A popular attempt in this direction is to estimate the winning chance of the player by so-called evaluation functions, and learning such evaluation functions utiliz- ing machine learning techniques became an important challenge of research in modern artificial intelligence. Artificial neural networks have been successfully used in many scientific and real world applications, for instance in pattern recognition, data mining, time series prediction. In recent years some attempts have been made to train artifi- cial neural networks for game playing tasks. Tesauro (1) has applied feedforward neural network models to play Backgammon where the artificial neural net was used together with reinforcement learning (RL) algorithms. Here in this paper a RL algorithm which is known as temporal difference learning has been inves- tigated for playing Tetris. The Tetris board is a grid of 10 columns and 20 rows of cells, which totals to 200 cells. Every cell can be in two possible states called

Read the paper · More papers on PaperTik