Distributed Multiagent Reinforcement Learning Approach for Multiserver Multiuser Task Offloading
He Zhang, Zijian Tian, Lanting Zeng, Limin Lu, Simengxu Qiao, Shifan Chen, Xinggao Liu · IEEE Internet of Things Journal · 2025
The industrial manufacturing industry requires user devices (UDs) to process massive data, leading to latency and energy consumption. To achieve low-latency and low-power task execution, we propose an industrial intelligent manufacturing system utilizing mobile edge computing. A dynamic computation offloading and resource allocation problem is formulated in multiserver multi-user scenarios to balance latency and energy cost. To solve this optimization problem, multi-agent deep reinforcement learning (MADRL) offers a theoretical framework. However, the widely used centralized training and decentralized execution (CTDE) scheme has two major drawbacks. First, it is unsuitable for scenarios without a central controller for global information collection. Second, it incurs significant communication overhead. Consequently, the decentralized training and execution (DTDE) scheme becomes necessary. However, DTDE introduces nonstationarity by treating other UDs as part of the environment, which leads to non-convergence. To address these issues, we design a distributed partial communication-based computation offloading and resource allocation algorithm (DPC-CORAA). This algorithm establishes a partial communication model based on the decentralized partially observable stochastic game (Dec-POSG) framework. It also incorporates the multi-agent deep deterministic policy gradient method under the DTDE scheme. The proposed method enables UDs to exchange information with neighbors to estimate the global decision, reducing communication cost. It also ensures theoretical convergence of this estimation to the real decision, serving as a local observation for independent strategy learning. Simulation results demonstrate that DPCCORAA achieves stable convergence, whereas the general DTDE scheme fails to converge. When contrasted with MADRL using CTDE, DPC-CORAA delivers superior performance in reducing latency and energy consumption, particularly in large-scale scenarios.