Multi-Task Resource Allocation and Task Offloading via Multi-Agent Deep Reinforcement Learning in Edge-Cloud system
Guoqing Tian, Xilong Wang, Xin Li, Xiaolin Qin · 2024
With the increasing complexity of deep neural network (DNN) models and the constrained computational and storage capabilities of user equipment (UE), the efficient inference of DNN models becomes a challenge. As an extension of cloud computing, edge computing was proposed to alleviate the pressure on cloud servers. However, the optimal offloading strategy and resource allocation for different types of DNN model tasks in Edge computing systems are still open problems. In this paper, we propose a DNN inference acceleration strategy based on deep reinforcement learning (DRL) for edge computing collaborative inference. Our approach aims to obtain the optimal DNN offloading strategy and resource allocation policy to achieve the lowest inference delay for each task request. We also consider the waiting time of tasks in resource-limited stations. Experimental demonstrate that our algorithm can decrease the average inference latency by as much as 62% than the compared to the Edge-Only algorithm.