A configuration of multi-agent reinforcement learning integrating prior knowledge
Hainan Tang, Hongjie Tang, Juntao Liu, Ziyun Rao, Yunshu Zhang, Xunhao Luo · 2024
Reinforcement learning aims to maximize the accumulated reward by interacting with an environment. However, with the increase in the number of agents, the problem called the curse of dimensionality arises because of the exponential growth of the joint state-action space. In this case, agents need to sample and explore as more as possible to update the policy gradient, in which the convergence speed is slow. So, it is difficult to apply in practice. To this end, we propose a configuration for multi-agent reinforcement learning algorithms to integrate prior knowledge. Based on the online reinforcement learning, prior knowledge is used to limit the action space in special states to reduce the sampling and exploration of the environment. Corresponding experiments suggest that the configuration can accelerate the learning process and help the agent learn the optimal policy faster.