A High-precision Quantization Method of Neural Networks for Concat Operator
Wei Dong Lu, Yang Chaojie, Yao Qin, Zhong Ma · 2023
The deployment of neural networks on the embedded devices is an important development trend in the future. Due to the limited resources on the embedded devices, the acceleration of neural networks is urgent and of great importance. Quantization, which converts floating-point neural networks into low-bit-width integer networks, is one of the most promising solutions for reducing the computing and storage cost of neural networks on the embedded devices. Concat is a common operator of neural networks. Since all the branches that are being concatenated should share the same quantization parameters, the quantization of the Concat operator results in a large loss of accuracy. It is the bottleneck problem that hinders the application of neural networks on embedded devices. In this paper, a high-precision quantization method for the Concat operator is proposed. The proposed method can replace the quantization of the Concat operator with the quantization of the Eltwise operator, which can effectively reduce the quantization accuracy loss. The results show that the accuracy of the proposed method can be improved up to 2.19% for classification, segmentation and small target detection applications.