Research on the Imperfect Information Game of Four-Player Mahjong Based on Mix-PPO

Jia-Yang Wang, Ming-Yan Wang, Wang Zeng, Zi-An Zhong · IEEE Transactions on Games · 2024

In recent years, the deep reinforcement learning method has performed well in many challenging tasks, includingGoandMOBAgames and other industrial fields.Mahjongis a popular game with imperfect information, but because of its large amount of hidden information and complex game rules, it is very challenging to solve its game intelligence decision problem and build artificial intelligence beyond the human level. To solve the aforementioned problems, this article proposes a feature encoding method and model training strategy for four players in ChineseMahjong. In addition, this article also innovatively proposes the Mix-PPO algorithm, which combines the advantages of the traditional PPO1 algorithm and the PPO2 algorithm, and compares the Mix-PPO algorithm with other algorithms, including the traditional proximal policy optimization algorithm, the deep-learning-related algorithm, and the game search tree algorithm. The experimental results show the validity of the feature coding and the Mix-PPO algorithm of Chinese four-playerMahjongin building the decision-making model of Chinese four-playerMahjong, as well as the validity of the coding method of the model'sMahjongfeature and the training strategy.

Read the paper · More papers on PaperTik