DNN Inference Acceleration and Reliable Task Offloading in Mobile-Edge Computing
Changke Wang, Xiaowen Huang, Wenqian Zhang, Guanglin Zhang · 2025
Most computation-intensive tasks are offloaded to high-performance edge servers, making the acceleration of deep neural networks (DNNs) a prominent research topic in mobile-edge computing networks. Deploying computationally demanding deep learning networks on resource-constrained mobile devices poses significant challenges. We first propose the overlapped partitioning DNNs, along with mobility-aware task offloading, focusing on analyzing the trade-offs between DNN inference acceleration and offloading reliability design insights. We then establish a collaborative DNN partitioning and offloading scenario, derive an acceleration and reliability model for DNN inference, and define the collaborative partitioning and offloading problem. Finally, we propose two algorithms: self-attention deterministic Policy Gradient (SAD) and optimal allocation algorithm (OAA). SAD globally optimizes task offloading strategies for mobile-edge DNNs, while OAA fine-tunes the strategy, optimizing inference acceleration and reliability between neighboring communication nodes. The simulation evaluations demonstrate the proposed scheme’s superiority.