A New Low-Bit Quantization Algorithm for Neural Networks

Xuanhong Wangl, Yuan Zhong, Jiawei Dong · 2023

Model quantization is one of the hot research topics in the field of model compression, which is important for reducing the model size and memory overhead. The current stage of model quantization methods usually uses straight-through estimators(STE) to train the network, which avoids the zero-gradient problem by directly replacing the derivative of the rounding function with the derivative of the constant function, but STE updates the values of the weights, in the dequantization propagation phase only propagate the same gradient. The discrete error between input and output is not taken into account, which will cause a large accuracy loss. It is not only difficult to achieve accuracy comparable to the full precision model, but also not stable enough to converge for large datasets with a high number of classifications. To address the shortcomings of floating-point gradient update and accuracy loss of low-bit quantization, we propose a new low-bit neural network quantization (LEQ) method. LEQ jointly optimizes the discrete error problem between input and output in backpropagation and the model stability problem, reduces the accuracy loss of the model during training, and uses symmetric cross-entropy loss function to enhance the robustness of the model, which has obvious advantages on several large public datasets. The accuracy of TOP-l can be 79.9,88.6,92.2,92.6, 98.7,98.4,99.1, and 99.1 when quantified to INT2, INT3, INT4, and INT8 on CIFAR-I0 and MNIST datasets, respectively.

Read the paper · More papers on PaperTik