Quantization Technology for Neural Network Models with Complex Branching Structures
Yang Yu, Wei Lu, Zhong Ma, Chaojie Yang, Yuejiao Wang · 2024
Model quantization is one of the means to address the difficulty of deploying models in resource-constrained environments. A quantization framework for integer-only inference across the entire network can better compress the resources required for computation and enhance the efficiency of model computation. Unlike mainstream model quantization frameworks that quantize by inserting quantization nodes into convolutional layers or fully connected layers, integer-only inference quantizes all operators in the model, including pooling and activation functions. However, in the case of multiple branches, there is a strong dependency relationship between the quantization parameters of model operators. Therefore, resolving the conflict of quantization parameters in complex branching structures is key to ensuring the accuracy of the quantization in the integer-only inference framework. This paper proposes a quantization technology for neural network models with complex branching structures, which resolves the conflict of quantization parameters in complex branching structures based on the topological relationships of the computation graph. This technology achieves a quantization loss of no more than 4% across models with different tasks and bit widths, demonstrating the effectiveness and practicality of this method.