ClipQ: Clipping Optimization for the Post-Training Quantization of Convolutional Neural Network
Yiming Chen, Hui Zhang, Chen Zhang, Yi Liu · Applied Sciences · 2025
In response to the issue that post-training quantization leads to performance degradation in mobile deployment, as well as the problem that the balanced consideration of quantization deviation by Clipping optimization techniques limits the improvement of quantization accuracy, this article proposes a novel clipping optimization method named ClipQ, which pays different attention to the parameters, aiming to preferentially reduce the quantization deviation of important parameters. The attention of the weight is positively related to its absolute value. Channel information entropy and principal component analysis are used to characterize the channel attention and spatial attention of activations, respectively. In addition, the particle swarm algorithm is applied in weight clipping to adjust the search step size and direction adaptively. ClipQ achieves high-precision quantization with very few calibration samples (<=50) and low time cost. Meanwhile, it does not bring extra computation, which is friendly to hardware. The experimental evaluation on image classification, semantic segmentation, and object detection shows that ClipQ outperforms other state-of-the-art clipping techniques, such as KL, ACIQ, and MSE. In 8-bit quantization, the average precision loss is 0.31% for image classification and 0.22% for object detection. More notably, it achieves almost lossless accuracy in semantic segmentation tasks.