Opponent Modeling by Expectation–Maximization and Sequence Prediction in Simplified Poker
Richard Mealing, Jonathan Shapiro · IEEE Transactions on Computational Intelligence and AI in Games · 2015
We consider the problem of learning an effective strategy online in a hidden information game against an opponent with a changing strategy. We want to model and exploit the opponent and make three proposals to do this; first, to infer its hidden information using an expectation-maximization (EM) algorithm; second, to predict its actions using a sequence prediction method; and third, to simulate games between our agent and our opponent model in-between games against the opponent. Our approach does not require knowledge outside the rules of the game, and does not assume that the opponent's strategy is stationary. Experiments in simplified poker games show that it increases the average payoff per game of a state-of-the-art no-regret learning algorithm.