AlphaZero applied to Jass

Andrin Bürli · Zenodo (CERN European Organization for Nuclear Research) · 2022

Despite recent successes of Artificial Intelligence applied to games, the performance ofcooperative agents in imperfect information games is still far from surpassing humans.In the traditional swiss card game of Jass four distinct parties play against each otherin two opposing teams. As an imperfect information game with competition as well ascooperation, Jass is a suitable candidate subject for research in this area. This project aims to evaluate if a purely reinforcement learning based method such asAlphaZero would be suitable for solving the game of Jass. However, since this methodis designed to work only on perfect information games, it cannot be used directly. For that reason, a new perfect information Jass is introduced where there is no hidden information. A fully asynchronous pipeline was designed to teach an AlphaZero based agent toplay the game of perfect information Jass. By means of a supervised dataset an optimized domain specific network architecture has been evaluated and trained. Using theimplemented pipeline in combination with this architecture, a variety of hyperparameters have been examined. The strongest configuration allowed an agent to achieve 0.75%average points againts the random opponent (APAR) and 0.48% average points againtsthe MCTS (100 simulations) based opponent (APAMCTS), therefore not beating thebaseline. After training, the agent is transferred to the imperfect information domainusing determinisations. In the new domain the best agent was using the network trainedon supervised data. It reached 0.72 APAR and at most 0.57 APAMCTS (20 simulations,25 determinizations) and therefore it is outperforming the MCTS baseline with statistical significance. Training an agent using self-play requires a lot of data and as a result also a lot of time.Considering the time and resource constraints of this project, the results do indicate thatthis approach might be suitable for the game of Jass. Nevertheless, further experimentswith more time and more computing power are required to prove this.

Read the paper · More papers on PaperTik