Research on Lightweight Compression Algorithms and Deployment Optimization for Large Language Models on Mobile Terminals
Yuan Li, Can Qu · 2025
With the development of artificial intelligence technology, Large Language Models (LLMs) have made significant achievements in the field of natural language processing. However, their large model size and computational complexity limit their application on resource constrained mobile devices. In this study, LLM is lightweighted through quantisation, pruning, knowledge distillation and other techniques to reduce model size and computational complexity. Combined with microservice architecture and cross platform compatibility optimization, it achieves rapid deployment and efficient operation of LLM on mobile devices. The experimental results of this study indicate that the lightweight compression algorithm effectively reduces the model size, improves deployment performance, and enhances user experience. It is of great significance to promote the widespread application of LLM on mobile intelligent devices.