A phased game algorithm combining deep reinforcement learning and UCT for Tibetan Jiu chess

Xiali Li, Yandong Chen, Yanyin Zhang, Bo Liu, Licheng Wu · 2023

The rules of the two phases of Tibetan Jiu chess, layout and battle, are very different, and using the same UCT search algorithm globally will result in a large overhead of time and storage space in the search process, so a phased game algorithm for Tibetan Jiu chess is proposed, with different strategies designed for the layout and battle phases, respectively. First, the layout phase uses a combination of Gaussian distribution and fast online estimation to improve the UCT algorithm, thus generating the optimal action selection scheme. Second, in order to take full advantage of reinforcement learning and deep learning, a neural network model with residual network structure is used in the battle phase to guide the search of Monte Carlo trees, and the default strategy is improved by "pruning" in the expansion step to improve the quality of the expanded nodes. The dataset is generated by self-play and used to train the neural network model to obtain the optimal model. It is verified through experiments that the phased gaming algorithm proposed in this study effectively reduces the process of blindly exploring the board state during the layout and battle phases of the UCT search algorithm, and improves the quality of the layout and the self-learning efficiency of the neural network model.

Read the paper · More papers on PaperTik