The MoE-Empowered Edge LLMS Deployment: Architecture, Challenges, and Opportunities
Ning Li, Song Guo, Tuo Zhang, Mu‐Qing Li, Zicong Hong, Qihua Zhou, Xin Yuan, Haijun Zhang · IEEE Communications Magazine · 2025
The powerfulness of LLMs indicates that deploying various LLMs with different scales and architectures on end, edge, and cloud to satisfy different requirements and adaptive heterogeneous hardware is the critical way to achieve ubiquitous intelligence for 6G. However, the massive parameters of LLMs poses significant challenges in deploying them on edge servers due to high computational and storage demands. Considering that the sparse activation in Mixture of Experts (MoE) is effective on scalable and dynamic allocation of computational and communications resources at the edge, this article proposes a novel MoE-empowered collaborative deployment framework for edge LLMs, denoted as CoEL. This framework fully leverages the properties of MoE architecture and encompasses three key aspects: model quantization, intra-server and inter-server cooperation, and token pruning and fusion. The CoEL begins with quantizing experts based on their importance and popularity, assigning different bit widths to different experts. Then, considering the heterogeneous resources of edge servers and model deployment requirements, a multi-dimensional collaborative deployment strategy is proposed. This strategy employs intra-server cooperation if the compressed model can be deployed on a single edge server; otherwise, it triggers inter-server cooperation and deploys experts across multiple edge servers distributed. Additionally, to minimize data transmission delays between servers, a token compression approach is applied. Finally, given the dynamic of network topology, resource status, and user requirements, the deployment strategies are regularly updated to maintain its relevance and effectiveness. This article also delineates the challenges and potential research directions for the deployment of edge LLMs.