A Simple Up-and-Down Weight Update Method for Tiny 8-bit Quantized CNN Training

Chanyung Kim, Eunchong Lee, Kyungho Kim, Sung‐Joon Jang, Sang-Seol Lee · 2024

Reducing model capacity of Convolution Neural Network (CNN) through low-bit quantization is challenging to maintain high accuracy of full-precision model. Quantization during training is particularly challenging due to a critical issue where the gradients during back-propagation are much smaller than weights, that they misleading in the update process. To solve this, various methods such as clipping, estimators, and adaptive scaling have been newly proposed. However, from a hardware perspective, adding computational costs to solve these problems is not desirable. To address this problem, we propose a method that maximizes the preservation of gradient-based ∆weight values typically lost in the weight update process. Our method converts and stores ∆weight values based on the weight quantization unit scale as a stack format. Using the Weight Stack (WS), weights can be updated simply by adding or subtracting the quotient without re-quantization of the weights. Comparing with full-precision models, the VGG-16 and ResNet-20 with 8-bit trained by our method achieves less than 0.1% accuracy drop on CIFAR-10.

Read the paper · More papers on PaperTik