DAPO: Mobility-Aware Joint Optimization of Model Partitioning and Task Offloading for Edge LLM Inference

Hao Feng, Huang Gan, Nian Zhou, Feng Zhang, Yuming Liu, Xiumin Zhou, Junchen Liu · Electronics · 2025

Deploying Large Language Models (LLMs) in edge environments faces two major challenges: (i) the conflict between limited device resources and high computational demands, and (ii) the dynamic impact of user mobility on model partitioning and task offloading decisions. To address these challenges, this paper proposes the Dynamic Adaptive Partitioning and Offloading (DAPO) framework, an intelligent solution for multi-user, multi-edge Mobile Edge Intelligence (MEI) systems. DAPO employs a Deep Deterministic Policy Gradient (DDPG) algorithm to jointly optimize the model partition point and the task offloading destination. By mapping continuous policy outputs onto valid discrete actions, DAPO efficiently addresses the high-dimensional hybrid action space and dynamically adapts to user mobility. Through extensive simulations, we demonstrate that DAPO outperforms baseline strategies and mainstream RL methods, achieving up to 27% lower latency and 18% lower energy consumption compared to PPO and A2C, while maintaining fast convergence and scalability in dynamic mobile environments.

Read the paper · More papers on PaperTik