Detecting and Tracing Multi-Strategic Agents with Opponent Modelling and Bayesian Policy Reuse

Hao Chen, Jian Xin Huang, Quan Liu, Chang Wang, Hanqiang Deng · 2020

In competitive multi-agent scenarios, the agents try to defeat their opponents by choosing the best response policies. However, non-stationary opponents make it difficult because they can also adapt to the evolved policies and behaviors of the agents. In this paper, we propose a novel Bayesian policy reuse approach for non-stationary opponents. It combines the learning of the best policy, the detection and prediction of the opponent policy, as well as the selection of the optimal response policy. We introduce an eXtended learning classifier system (XCS) for multi-agent reinforcement learning algorithm in Markov games. Besides, we incorporate the opponent models for opponent policy identification and prediction. Furthermore, we propose a novel online policy reuse technique which can accurately and quickly trace the opponents' policies in tasks with different rewards. We demonstrate the performance of the proposed approach by comparing it with state-of-art existing algorithms in competitive Markov games.

Read the paper · More papers on PaperTik