Clipping-based Neural Network Post Training Quantization for Object Detection
Cui Liqun, Lei Hu · 2023
Convolutional Neural Networks (CNN) have achieved amazing breakthroughs in computer vision, speech recognition, and other fields. However, with the development of deep learning, the number of weights in the neural network is increasing, consuming a lot of storage and computing resources. The high storage cost and computational complexity severely restrict the deployment of deep learning on embedded mobile devices. Therefore, the compression and acceleration of convolutional neural networks becomes particularly important. Therefore, in order to compress the size of the target detection model and speed up the calculation, this paper proposes a post-training quantization method for clip limit operation, limits are performed on the input data stream and weight parameters during training, then uses the parameter quantization method to quantize the model parameters from 32-bit floating point to 8-bit Integer. Under the premise of no impact on the performance of the algorithm, the model can be effectively compressed and the calculation speed can be improved. On the public data set PASCAL VOC2012, the performance of this algorithm is evaluated by two classic target detection algorithms, Faster R-CNN and YOLO V3. The experimental results show that compared with the original model, it saves 75% of the storage space and improves the calculation speed 2 times faster; FPS is also improved compared to standard quantization.