Multi‐Agent Reinforcement Learning With Deep Networks for Diverse Q‐Vectors

Zhenglong Luo, Zhiyong Chen, Shijian Liu, James Stuart Welsh · Electronics Letters · 2025

ABSTRACT In multi‐agent reinforcement learning (MARL) tasks, the state‐action value, commonly referred to as the ‐value, can vary among agents because of their individual rewards, resulting in a ‐vector. Determining an optimal policy is challenging, as it involves more than just maximizing a single ‐value. Various optimal policies, such as a Nash equilibrium, have been studied in this context. Algorithms like Nash Q‐learning and Nash Actor‐Critic have shown effectiveness in these scenarios. This paper extends this research by proposing a deep Q‐networks algorithm capable of learning various ‐vectors using Max, Nash, and Maximin strategies. We validate the effectiveness of our approach in a dual‐arm robotic environment, a representative human cyber‐physical systems (HCPS) scenario, where two robotic arms collaborate to lift a pot or hand over a hammer to each other. This setting highlights how incorporating MARL into HCPS can address real‐world complexities such as physical constraints, communication overhead, and dynamic interactions among multiple agents.

Read the paper · More papers on PaperTik