Multi‐Agent Reinforcement Learning With Deep Networks for Diverse Q‐Vectors
Zhenglong Luo, Zhiyong Chen, Shijian Liu, James Stuart Welsh · Electronics Letters · 2025
ABSTRACT In multi‐agent reinforcement learning (MARL) tasks, the state‐action value, commonly referred to as the ‐value, can vary among agents because of their individual rewards, resulting in a ‐vector. Determining an optimal policy is challenging, as it involves more than just maximizing a single ‐value. Various optimal policies, such as a Nash equilibrium, have been studied in this context. Algorithms like Nash Q‐learning and Nash Actor‐Critic have shown effectiveness in these scenarios. This paper extends this research by proposing a deep Q‐networks algorithm capable of learning various ‐vectors using Max, Nash, and Maximin strategies. We validate the effectiveness of our approach in a dual‐arm robotic environment, a representative human cyber‐physical systems (HCPS) scenario, where two robotic arms collaborate to lift a pot or hand over a hammer to each other. This setting highlights how incorporating MARL into HCPS can address real‐world complexities such as physical constraints, communication overhead, and dynamic interactions among multiple agents.