Clipped Quantization Aware Training for Hardware Friendly Implementation of Image Classification Networks
Kyungchul Lee, Jongsun Park · 2022 19th International SoC Design Conference (ISOCC) · 2022
Although deep neural networks (DNNs) show excellent performance in the image processing field, a massive amount of computation and memory access makes it difficult to deploy DNNs on mobile devices. To reduce the computation and the memory access, quantization shows great results. However, it is challenging to quantize the network into low bit-width without significant accuracy degradation. In this paper, we propose Clipped Quantization Aware Training (CQAT) to decrease the accuracy drop during the low bit-width quantization aware training. In the proposed CQAT, the original values are first quantized into 8-bit and then the quantized values are clipped so that only 4-bit for activations and 5-bit for weights are used. With the proposed technique, ResNet-18 for the CIFAR-100 dataset using 5-bit weight and 4-bit activation shows an accuracy drop of only 0.96% compared to the network using full precision weight and activation.