An intelligent decision making method based on Bayesian policy reuse framework
Yixin Zhang, Qi Han, Li Zhu · 2022 3rd International Conference on Electronic Communication and Artificial Intelligence (IWECAI) · 2022
Transfer learning plays an important role in multi-agent reinforcement learning. Bayesian policy reuse(BPR) algorithm is an excellent transfer learning algorithm and has been successfully used in repeated stochastic games. BPR can quickly detect the opponents and select the optimal policy. BPR's main limitation is the need of an offline learning phase where the models (priori knowledge) can be obtained. However, BPR may perform poorly when there is a lack of prior knowledge or the prior knowledge is inaccurate. To solve this problem, we propose an algorithm that combine the no-regret with BPR. Our approach can be exploited to improve performance when the priori knowledge is accurate, and to achieve less poor performance when the priori knowledge is inaccurate. Our approach has robust theoretical guarantees and is validated on the repeated stochastic games.