Designing Quantizers for Low-Precision Post-Training Quantization: A Standard Pipeline Approach for CNNs
Yang Xiao, Fei Chao · 2024
Quantization is a widely adopted technique for enhancing the inference performance of neural networks by exploiting hardware support for low-precision algorithms. It offers significant advantages such as higher throughput, reduced memory traffic, and decreased storage requirements. Neural networks have demonstrated robustness to quantization, allowing them to be quantized to lower bit widths with minimal impact on accuracy. However, the quantization process often leads to a loss in precision, especially when performed without retraining. Addressing this issue, this paper focuses on the design of quantizers for low-precision post-training quantization tasks, aiming to minimize accuracy loss without the need for access to the full training dataset or extensive computational resources. We propose a novel approach that introduces outlier clipping for weights and designs bias correction schemes for both weights and activations. Our methodology emphasizes the development of a standard pipeline for low-precision post-training quantization, which balances simplicity and performance. Through comprehensive experiments, we demonstrate that our proposed pipeline achieves superior accuracy compared to many existing post-training quantization methods. This work contributes to the ongoing efforts in the field by providing a practical and efficient solution for quantizing convolutional neural networks (CNNs) in low-precision scenarios.