ATA-MAOPT: Multi-Agent Online Policy Transfer Using Attention Mechanism With Time Abstraction

Chao Wang, Xiangbin Zhu · IEEE Access · 2024

Multi-agent deep reinforcement learning inherits one of the shortcomings of reinforcement learning, which is that a very large number of experimental episodes need to be performed to train an agent model to learn the optimal policy. Transfer learning is a typical method to accelerate the training process in machine learning field. In cooperative multi-agent deep reinforcement learning, we can reasonably design transfer learning methods to accelerate the learning of homogeneous agents. Transfer learning in the field of multi-agent systems is different from that in single agent environment, as in the systems there are multiple agents learning at the same time, online policy transfer methods can be used among these agents. In this paper, we propose an online policy transfer algorithm based on attention mechanism and time abstraction for multi-agent deep reinforcement learning (ATA-MAOPT). The algorithm employs the option-critic framework, which is a typical temporal abstraction mechanism and utilizes the option mechanism to select the source policy for policy transfer, and at the same time, the algorithm introduces the attention mechanism into the algorithm to make these options focus more on important features in observation spaces, making the source policy more accurate and improving the efficiency of policy transfer. The experimental results demonstrate that the algorithm can effectively speed up the learning process of the agents and reduce the computational costs.

Read the paper · More papers on PaperTik