Analysis of Neural Network Accuracy Degradation due to Uniform Weight Quantization of One or More Layers
Jelena Nikolić, Stefan S. Tomić, Zoran Perić, Danijela Aleksić · 2022 57th International Scientific Conference on Information, Communication and Energy Systems and Technologies (ICEST) · 2022
In this paper, our goal is to examine which layer of the neural network (NN) is the most sensitive to the two-bit uniform weight quantization and results in the highest accuracy degradation of our trained NN model. We perform four NN accuracy analyses to qualify the efficiency of uniform two-bit quantization for each NN layer. Layer quantization importance ranking provided conclusion that the most important layer is the first layer, with the largest number of weights, while the last layer, although containing the smallest number of weights, is the second important one. This conclusion could be especially useful for preliminary anticipation in larger NN architectures with mixed-precision quantization.