A Survey On Neural Network Quantization
Jiawei Yang, Zhongbo Li, Zeqin Feng, Yongqiang Xie · 2025
In recent years, with the extensive implementation of neural network models in fields such as computer vision and natural language processing, the issue of substantial model parameters and significant computational resource utilization has come to the fore. Neural network quantization techniques have emerged as a primary solution to address the constraints imposed by model deployment resources by reducing the numerical precision both of model parameters (include weights, activation values and gradients). This paper undertakes a systematic exploration of quantization methods employed in traditional neural networks, including Convolutional Neural Networks and Recurrent Neural Networks, as well as neural networks based on the Transformer architecture. It delves into the technical distinctions between these methods and examines the challenges and limitations in their application. Additionally, it discusses the future directions and trends in this field.