A Survey of AI Inference Technologies for On-Device Systems
Wenzhu Wang, Ke Li, Bin Ji, Xiaodong Liu, Jie Yu, Qingbo Wu · IEEE Internet of Things Journal · 2025
In recent years, artificial intelligence(AI) technologies represented by foundation models have experienced rapid development. Concurrently, On-device AI inference has become the primary approach for intelligent technology applications, offering advantages such as low latency, high security, and personalization. However, due to the limited resources of on-device systems, on-device AI inference faces new challenges, including improving computational efficiency, optimizing task parallelism, and model optimization. This survey addresses these challenges from a software and algorithmic perspective, focusing on three key areas: Operator Computation: Explores methods to accelerate matrix multiplication and convolution, as well as techniques like operator fusion and vectorized computation. Task Inference: Analyzes heterogeneous and distributed computing, memory allocation, and energy-efficient tuning to improve the parallel execution and energy efficiency of inference tasks. AI Models: Covers model compression, lookup table quantization, and model architecture design to reduce computational complexity and storage requirements. By analyzing these areas, the survey aims to improve inference speed, reduce resource dependency, and provide insights into the future trends of on-device AI technology.