Maximum Entropy Reinforcement Learning in Two-Player Perfect Information Games

Taichi Nakayashiki, Tomoyuki Kaneko · 2021 IEEE Symposium Series on Computational Intelligence (SSCI) · 2021

This paper studies maximum entropy reinforcement learning in two-player adversarial games. Reinforcement learning (RL) has been effective for mastering such games since Alp-haZero carefully integrates Monte-Carlo tree search and self-play. Maximum entropy RL is a promising enhancement that maximize the entropy of a policy (or strategy in games) in addition to cumulative rewards (or wins in games), yielding effective methods such as soft Q-learning or soft actor-critic. The entropy objective improves the robustness of learned strategies by keeping (reducing) alternative choices on our (the opponent's) side when applied to games. Moreover, we show that maximum entropy RL also improves the stability and efficiency of learning by reducing the variance in learning targets of neural networks, which are typically {1, 0, -1} for win, draw, or loss in AlphaZero. Being augmented with the entropy and integrated through a game tree, maximum entropy RL yields soft state values taking intermediate real values, especially near zero in positions near opening positions. We qualitatively show these benefits in terms of an analysis of converged values done by extended retrograde analysis for dobutsu shogi and empirical learning efficiency for RL agents in self-play.

Read the paper · More papers on PaperTik