Joint Inference Offloading and Model Caching for Small and Large Language Model Collaboration

Xinyi Xu, Gang Feng, Yi-Jing Liu, Shuang Jian Qin, Jian Wang, Yunxiang Wang · IEEE Transactions on Mobile Computing · 2025

Large Language Models (LLMs), with advanced content creation and inference capabilities, can provide immersive intelligent services to users in mobile edge networks. However, the increasing demand for real-time artificial intelligence (AI) applications aggravates the limitations of cloud-based LLMs due to the long response time. Meanwhile, Small Language Models (SLMs), which are cost-effective and locally deployable for terminal devices, can serve as an efficient supplement to LLMs for performing latency-sensitive tasks with lower generalization capability. Due to the resource constraints of edge networks and the diverse requirements of user tasks, it is critical to design an inference framework that effectively coordinates the deployment and collaboration of LLMs and SLMs. In this paper, we propose an LLM-SLM collaborative inference (LSCI) scheme under a mobile edge computing (MEC) architecture, which jointly decides where to cache models and how to offload inference tasks to balance latency, accuracy, and resource costs. To optimize inference performance subject to resource constraints, we jointly solve the inference task offloading and model caching problem in LSCI scheme. Specifically, we employ deep reinforcement learning (DRL) to select highly popular SLMs to be cached on the edge server, and distributed belief propagation technique to solve the associated inference task offloading issue. Numerical results show that the proposed LSCI scheme can achieve significant performance gain in terms of inference performance when compared with a number of baseline solutions.

Read the paper · More papers on PaperTik