MDou: Accelerating DouDiZhu Self-Play Learning Using Monte-Carlo Method With Minimum Split Pruning and a Single Q-Network

Qian Luo, Tien-Ping Tan, Yi Su, Zhanggen Jin · IEEE Transactions on Games · 2022

Artificial intelligence (AI) has demonstrated outstanding performance in some perfect- and imperfect-information games, such asGo, Atari, and Texas Hold'em. Even though AI is successful in these games with small action spaces, it does not play well in large-scale multiplayer, imperfect-information games likeDouDiZhu. DouZero, a DouDizhu AI system, has recently been proposed and beaten all the existing DouDizhu AI programs. This article introduces minimum split pruning (MSP) and a singleQ-network to accelerate the training of DouZero, called MDou. Our experiments show that MDou improved through self-play using limited computational ability (only a 4-core CPU and 1 GPU) and less learning time (30 days), while achieving comparable performance to DouZero.

Read the paper · More papers on PaperTik