Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference

Le Chen, Dahu Feng, Erhu Feng, Yingrui Wang, Rong Juan Zhao, Yubin Xia, Pinjie Xu, Haibo Chen · 2025

With the rapid advancement of artificial intelligence technologies such as ChatGPT, AI agents, and video generation, contemporary mobile systems have begun integrating these AI capabilities on local devices to enhance privacy and reduce response latency. To meet the computational demands of AI tasks, current mobile SoCs are equipped with diverse AI accelerators, including GPUs and Neural Processing Units (NPUs). However, there has not been a comprehensive characterization of these heterogeneous processors, and existing designs typically only leverage a single AI accelerator for LLM inference, leading to suboptimal use of computational resources and memory bandwidth.

Read the paper · More papers on PaperTik