Edge vs Cloud: How Do We Balance Cost, Latency, and Quality for Large Language Models Over 5G Networks?

Minsu Kim, Pinyarash Pinyoanuntapong, Bongho Kim, Walid Saad, Doru Calin · 2025

Large language models (LLMs) can perform a plethora of tasks, however, they often require cloud servers for deployment due to their computing cost and size. Meanwhile, small and cost-effective LLMs can be deployed on edge devices (e.g., mobile devices), but they often exhibit lower response quality than larger models. In this paper, a measurement-driven training framework is proposed for a device-side artificial intelligence (AI)-enabled router that selects the best LLM in terms of cost, latency, and performance. In the considered framework, a mobile device uses its local LLM (sub-billion LLM), while having access to a bigger server LLM (GPT-4) hosted on a 5G core network. The mobile device has a router that sends an input prompt to the local or server LLM to optimize the cost, latency, and performance of output LLM responses. To train the router, the dataset is constructed by measuring the cost, latency, and performance of the server LLM and the local LLM on a mobile device. The structural causal model (SCM) of the measured dataset is identified. To further improve performance, a causality-driven data augmentation method is also proposed based on the discovered SCM. Real-world experimental results show that the proposed framework can improve the cost and latency by 50% and 31%, respectively, with only a 2.13% performance loss compared to a baseline that only uses the server LLM.

Read the paper · More papers on PaperTik