Time-delayed Data Transmission in Heterogeneous Multi-agent Deep Reinforcement Learning System

Elhami Fard, Rastko R. Šelmić · 2022 30th Mediterranean Conference on Control and Automation (MED) · 2022

This paper studies the data transmission between agents of a multi-agent, deep reinforcement learning (MADRL) system (leaderless and leader-follower) using the deep Q-network (DQN) algorithm. The structure of the MADRL system consists of various clusters of agents. The agents in a cluster have the same architectures. The DQN architecture is used to present the first cluster’s agents structure. The other clusters, including various architectures, are considered as the environment of the first cluster’s deep reinforcement learning (DRL) agent. The goal of each static agent is to transfer data with the maximum average reward. We consider two novel observations in data transmission termed on-time and time-delay. The two proposed observations are considered when the data transmission channel is idle and the data is transmitted on-time or time-delayed. Moreover, by considering the distance between the neighboring agents, we present a novel immediate reward function by appending a distance-based reward to the previously utilized reward. We have rigorously shown which system (on-time or time-delayed) has a superior performance based on the DQN loss and team reward for the entire team of agents. The claims have been proven theoretically, and the simulation confirms theoretical findings.

Read the paper · More papers on PaperTik