Birds in Cages: Edge Inference Allocation for Distributed LLM Deployment
Jiahao Zhu, Lu Zhao, Fu Xiao, Lingjie Duan · 2025
The distributed deployment of Large Language Models (LLM) on edge servers close to users has unlocked the service provider's potential to deliver low-latency inference. To obtain more benefit by serving more resource-demanding inference tasks based on resource-limited edge servers, it is critical for the service provider to allocate inference to suitable edge servers. Three new challenges hinder existing approaches from being implemented: the distributed LLM inference requires edge servers to collaborate following a novel workflow different from other tasks; the generative nature of LLM incurs uncertainty in task resource occupations; considering the heterogeneity in users' latency requirements and service benefit, merely minimizing the total user-perceived latency can not maximize the benefit. In this paper, we make the first attempt to study the edge inference allocation problem for distributed LLM deployment while conquering these challenges. Specifically, we propose a collaborative workflow for edge servers to conduct distributed LLM inferences. Then, we estimate the resource occupations by employing Exact Conic Reformulation (ECR). Based on this, with the objective of maximizing the total service benefit, we formulate the inference allocation problem as a binary integer programming problem which is NP-hard. An approximate algorithm is proposed to find approximate solutions efficiently. Extensive experiments based on a real-world dataset demonstrate the performance of our approach.