Progress of neural network quantization algorithm

Zhenzhe Wang, Yuxiang Yao, Yucheng Zhou · 2021 International Conference on Electronic Information Engineering and Computer Science (EIECS) · 2021

Deep neural networks have achieved great success in various fields of artificial intelligence nowadays. However, the inference process of deep learning models requires a huge number of parameters and calculations. Therefore, how to deploy on mobile devices has become a problem. Any bit in the embedded hardware design is very important to the neural network. Extremely low bits are used to represent neural networks to reduce the consumption of resources such as memory, but it usually leads to a serious loss of accuracy. So how to train an extremely low-level neural network with high accuracy is very important. Most of the existing neural network quantization methods are based on low-level weight quantization or low-level activation quantization. This article reviews various neural network quantization methods commonly used in recent years and compares their effects. Compared with previous reviews, the quantization techniques introduced in this article are more comprehensive, and the classification of various methods is clearer. After comparison, the greatest difficulty of the quantization method is the large loss of accuracy on large-scale data sets. Compared with the full-precision model, the quantization model of 8bit and above can basically achieve accuracy without any loss.

Read the paper · More papers on PaperTik