Bayesian Opponent Exploitation by Inferring the Opponent’s Policy Selection Pattern
Kuei-Tso Lee, Sheng‐Jyh Wang · 2022
In a multi-agent competitive domain, the agent needs to anticipate the opponent’s behavior and select a suitable policy to exploit the opponent. In this work, based on the BPR (Bayesian Policy Reuse) framework, we further assume the opponent may determine its policy depending on its previous observation. To deal with opponents of this kind, we discuss three different approaches for the agent, including learning from scratch, reasoning from experience, and reasoning accompanied by learning. The “reasoning accompanied by learning” approach turns out to be the most favorable method, in which the agent executes an iterative process that alternates between “updating the belief of each pre-collected model” and “progressively learning the opponent’s policy selection pattern” based on the observed data. In our experiments, we simulate a simplified batter vs. pitcher game. The experimental results show that the “reasoning accompanied by learning” approach does receive a larger averaged utility value than the learn-from-scratch approach and the reason-from-experience approach.