DeepHoldem: An Efficient End-to-End Texas Hold'em Artificial Intelligence Fusion of Algorithmic Game Theory and Game Information

Ke Wang, Dongdong Bai, Qibin Zhou · 2022

Texas Hold'em has received considerable attention from researchers as a typical representative of imperfect information games. In the last 5 years, Texas Hold'em artificial intelligence (AI), which adopts algorithmic game theory as its core technology, has achieved many breakthroughs, with Libratus, DeepStack, Pluribus, and other Texas Hold'em AIs achieving the ability to beat top human players in heads-up and multiplayer no-limit Texas Hold'em. In the past 2 years, researchers have begun exploring the use of neural networks to build decision-making models to reduce the dependence of model training on expert knowledge. AlphaHoldem is an essential representative of these neural networks, beating Slumbot through end-to-end neural networks. However, AlphaHoldem does not fully consider game rules and other game information, and thus, the model's training relies on a large number of sampling and massive samples, making its training process considerably complicated. In this study, we propose DeepHoldem, an efficient end-to-end Texas Hold'em AI that combines algorithmic game theory and game information. DeepHoldem uses reinforcement learning to simulate the self-play process of the counterfactual regret minimization algorithm. DeepHoldem can efficiently converge to the Nash equilibrium by considering player and opponent hand information when encoding the game state and incorporating the attention module into the network structure to achieve the correlation extraction of game action sequences, and thus, fully utilize game state information. In addition, DeepHoldem incorporates the concept of supervised learning into the training process, dramatically improving the training efficiency of the model and expanding its application scope. Experimental results show that after about 4 days of training, DeepHoldem beats Slumbot at a level of 2.1 mbb/h, and the single decision time is less than 0.1 seconds.

Read the paper · More papers on PaperTik