Reducing Neural Networks Quantization Error Through the Change of Feature Correlation

Ziheng Zhu, Shuzheng Wang, Haoyu Li, Xinyuan Tian · 2024

The quantization technique of neural networks can achieve a compressed representation of models by reducing weights and activations data bitwidth, accelerating the inference process, and minimizing models' consumption of memory occupation and calculation bandwidth. However, existing quantization methods only focus on the change in weight distribution, unable to directly form a feedback relationship between network feature extraction capability and quantization function, which makes the calculation of quantization error and the adjustment of quantization parameters not reasonable, at last leads to unstable convergence and reduced accuracy. We propose a quantization-aware training (QAT) method based on feature loss, pay more attention to the impact of quantization on feature space changes, and compute the topology changes within and between sample classes. The scheme is verified to reduce the performance gap between quantized and floating-point models.

Read the paper · More papers on PaperTik