Hybrid Actor–Critic for Physically Heterogeneous Multiagent Reinforcement Learning
Tianyi Hu, Zhiqiang Pu, Xiaolin Ai, Tenghai Qiu, Yanyan Liang, Jianqiang Yi · IEEE Transactions on Cognitive and Developmental Systems · 2025
This paper focuses on cooperative policy learning for physically heterogeneous multi-agent system (PHet-MAS), where agents have different observation spaces, action spaces and local state transitions. Due to the various input-output structures of agents’ policies in PHet-MAS, it’s difficult to employ parameter sharing techniques for sample efficiency. Moreover, a totally heterogeneous policy design impedes agents from utilizing the training experience of their companions, and increases the risk of environmental non-stationarity. To address the above issues, we proposehybrid heterogeneous actor-critic(HHAC), a method for the policy learning of PHet-MAS. The framework of HHAC consists of a hybrid actor and a hybrid critic, both containing globally shared and locally shared modules. The locally shared modules can be customized according to the actual physical properties of agents, while the globally shared modules can help extract and utilize the common information among agents. In the hybrid critic, a behavioral intention module is designed to alleviate the environmental non-stationary issue caused by evolving heterogeneous policies. Finally, a hybrid network training method is developed to address challenges in sample construction and training stability of hybrid networks. As evidenced by experimental results, HHAC exhibits superior performance enhancements over baseline approaches, and can facilitate PHet-MAS in learning sophisticated and instructive policies.