An Investigation of the Behaviours of Machine Learning Agents Used in the Game of Go
Mingyue Zhang, Xiao–Yi Zhang, Paolo Arcaini, Fuyuki Ishikawa · 2023
In last years, Machine Learning (ML) techniques have been successfully applied in various domains. Notably, in 2015, AlphaGo, an artificial Go game player (Go agent), based on deep reinforcement learning and Monte Carlo Tree Search, won by 4:1 against the renowned player Lee Sedol, which demonstrated the power of ML techniques. Although Go is a discrete and zero-sum game, compared with other board games it is much more complex. First, it has around 10172positions (i.e., placement of stones) which makes it impossible to find a perfect strategy to win the game. Secondly, playing the game requires not only the calculation of future moves but also strategies such as judgement of thickness, the balance between outside influence and territory, etc. The Go community is interested in understanding why and when Go agents are successful, in order to better understand the Go game itself. However, Go agents are inherently unexplainable, which makes it difficult to understand their strategies. In this paper, we make a first step towards this direction. We start from the observation that explaining a single agent is too difficult, while more insights can be obtained by comparing the strategies of multiple agents. Based on this, we perform an empirical study in which we compare different Go agents having different skill levels, measured by their Elo scores. We generate several different Go positions and we play all the agents from them. Then, we compare their decisions to understand when they played the same move and when they played differently. Moreover, we assume that, even if two agents play a different move, their strategies can be considered similar. The strategy of a Go agent is specified as a ranking of moves from which it usually picks the top one as move to play next. Therefore, to assess to what extent the strategies of different ML agents are similar, we compare the rankings they produce using the similarity metric Rank-Biased Overlap (RBO).