Neural Network Quantization Based on Model Equivalence

Huiqu Yang, Jian Xu, Guowei Yang, Ming Zhang, Hong Qin · 2022

With the rapid development of deep learning the size of neural networks has become increasingly large and it has become more and more difficult to deploy large neural networks on lightweight devices as well as small devices. Neural network model compression has become a hot research topic as a major approach to solving the above problems. One of the common compression techniques is model quantization but current quantization techniques ignore the equivalence between quantization models and floating-point models and simply use the performance metrics of floating-point models as the performance metrics of quantization models. For this reason this paper proposes a new evaluation criterion of model equivalence and proposes a neural network quantization algorithm based on the model equivalence criterion. According to the experimental findings the algorithm-quantized ResNet50 ensures no decrease in accuracy while improving the equivalence rate on the CIFAR100 dataset by 0.97% as compared to the conventional approach under eight-bit precision. In contrast the accuracy of the algorithm-quantized ResNet50 on the ImageNet dataset only decreases by 0.7%

Read the paper · More papers on PaperTik