Adaptive Model Partitioning and Pruning for Collaborative DNN Inference in Mobile Edge-Cloud Computing Networks

Hui Li, Xiuhua Li, Qilin Fan, Qiang He, Xiaofei Wang, Victor C. M. Leung · IEEE Transactions on Mobile Computing · 2025

Deep neural network (DNN) model partitioning and pruning have proven to be effective methods for enhancing resource efficiency and reducing inference delay by strategically allocating DNN workloads across heterogeneous edge and cloud infrastructures. Nevertheless, the heterogeneous nature of resources complicates the deployment of DNN in mobile edge-cloud computing (MEC) networks. In this paper, we present an innovative framework for collaborative DNN inference in MEC networks by integrating fine-grained model partitioning and magnitude-based pruning. However, the joint model partitioning and pruning policy presents significant challenges due to the inherently coupled and mutually influential nature. To address it, we adopt Long Short-Term Memory (LSTM) networks as action generation controllers to generate discrete actions for model partitioning and pruning alternately. After that, we adopt the policy gradient algorithms to optimize the LSTM-generated actions with a moving average according to the Monte Carlo estimate. By directly optimizing the policy function, the proposed framework enhances the efficiency and stability of action space exploration, yielding faster convergence and improved inference performance. Experimental results on standard datasets indicate that the proposed framework outperforms state-of-the-art approaches, achieving an 8.247% increase in system reward and an average reduction of 27.313% in total delay within the considered MEC networks.

Read the paper · More papers on PaperTik