Channel-wise quantization without accuracy degradation using Δloss analysis

Kazuki Okado, Kengo Matsumoto, Atsuki Inoue, Hiroshi Kawaguchi, Yasufumi Sakai · 2022

Recent studies have pointed out that the effect of quantization of convolutional neural networks on accuracy varies from layer to layer. For this reason, partial quantization or mixed-precision quantization on a layer basis have been considered for quantization. However, the layer quantization has a large impact on accuracy because its granularity is large; so, it generally requires retraining of the network, which incurs high computational cost. In this study, we proposed a new search algorithm for partial quantization, which aims to derive practical combinations of quantized channels without retraining. The proposed method successfully quantizes 83.3% of the parameters without degrading the accuracy in ResNet18 4-bit quantization. In addition, the proposed method succeeded in compressing 80.8% of the parameters in ResNet34 without degrading the accuracy.

Read the paper · More papers on PaperTik