Fast On-device LLM Inference with NPUs
Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu, Gang L. Huang, Mengwei Xu, Xuanzhe Liu · 2025
On-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest. However, even mobile-sized LLMs (e.g., Gemma-2B) encounter unacceptably high inference latency, often bottlenecked by the prefill stage in tasks like screen UI understanding.