Model Caching and Application Offloading for Mobile Edge Intelligence Network With Learning-and-Optimization Approach
Ziyu Peng, Yu Qiu, Gaocai Wang · IEEE Transactions on Services Computing · 2025
Mobile Edge Intelligence is a promising computing paradigm for mobile users to access Artificial Intelligence (AI) services. It seamlessly integrates AI online inference processes with Mobile Edge Computing (MEC), delivering low-latency services through application offloading. However, previous works often overlooked the need to pre-deploy relevant AI models on the edge server and disregarded the impact of model deployment on service performance. Furthermore, even with proper model deployment, service efficiency still depends heavily on the resource allocation strategies of the edge server. To this end, we propose a hybrid Deep Reinforcement Learning (DRL) approach, termed SA2CNN, which jointly optimizes discrete decisions on AI model caching and application offloading, along with continuous allocation of bandwidth and frequency resources, to minimize user energy cost and task latency. We first formulate the aforementioned challenges into a Markov decision process, then decompose it into two low-complexity sub-problems: the Discrete Destination Selection (DDS) problem and the Continuous Resources Allocation (CRA) problem. DRL is responsible for outputting the DDS actions, while CRA sub-problem results are solved by convex optimization. Furthermore, we decouple the caching and offloading decisions across time slots to eliminate the impact of model deployment on task performance, and employ a deep convolutional neural network to effectively learn the underlying temporal dependencies. Simulations show that our approach reduces service cost by 52.1% to 68.1% compared to the baselines, exhibiting significant performance enhancements.