MobiSplit: Mobility-Aware Inference Partitioning and Offloading for Efficient Edge Intelligence

Peng Wang, Wen Yao Sun, Yi Xin Yang, Dusit Tao Niyato, Dapeng Oliver Wu · IEEE Transactions on Mobile Computing · 2025

Edge intelligence enhances the computational capabilities of resource-limited devices by offloading inference tasks to edge servers. Traditional methods either execute the entire model on the device, resulting in slow inference, or fully offload it to the server, incurring communication delays and privacy risks due to raw data transmission. Model partitioning addresses these challenges by splitting the model for execution on both the device and edge server, transmitting only intermediate inference results. However, current model partitioning methods lack consideration of device mobility, resulting in reduced inference efficiency and task interruptions. To address these limitations, we introduce MobiSplit, a novel mobility-aware framework that dynamically partitions inference models between resource-constrained devices and edge servers. MobiSplit adapts to real-time device mobility, fluctuating network conditions, and computational constraints to minimize inference latency and energy consumption while ensuring robust task execution. Additionally, we propose a distributed auction-based algorithm that empowers edge devices to autonomously determine optimal partitioning and offloading strategies in a scalable and adaptive manner. Extensive simulations demonstrate that MobiSplit enhances inference efficiency, achieving a 60% latency reduction and a 20% energy consumption decrease compared to the best-performing baseline across diverse edge scenarios.

Read the paper · More papers on PaperTik