Towards Expert Models Deployment Cost Optimization in Edge Computing Networks
Jiaqi Ren, Chao Wang, Yihan Zhong, Shaohua Cao, Danyang Zheng, Xiaojun Cao · 2025
With the widespread adoption of large language models (LLMs) like GPT, user experiences in various interactive applications have significantly improved. However, reports from OpenAI highlight that GPT clients are now facing high response delays and frequent interruptions, particularly during peak usage hours, due to limited computation resources. This challenge is expected to escalate as machines are interacting with GPT models at higher frequencies, with greater data volumes, and over longer lifecycles. A promising solution is to deploy LLMs across edge networks to efficiently distribute the huge resource demands. This work presents the very first efforts at exploring how to cost-effectively deploy expert models from a mixture of experts (MoE) LLM within edge networks. We introduce the expert models deployment in edge networks (EMD-EN) problem, focusing on optimizing deployment costs. To address this, we propose a novel least cost gain (LCG) measure for selecting appropriate physical nodes to host expert models and present a corresponding LCG-based expert models deployment (LCGEMD) algorithm. Extensive simulations show that our approach outperforms the benchmarks by an average of 17.31% and 36.98% in terms of deployment cost reduction.