A Novel Quantization with Channel Separation and Clustering for Deep Neural Networks
Fangzhou He, Ke Ding, Zhicheng Dong, Jiajun Wang, Mingzhe Chen, Jie Li · 2024
The significant computational demands and memory needs of deep neural networks (DNNs) pose challenges for deploying them in edge learning systems, such as vehicular networks and internet-of-things applications. While model quantization is proposed to improve the efficiency for edge inference of DNNs, existing quantization schemes have large memory overhead, and need complex processes and long hours of training. Although a subset of quantization methods exists that are computationally simple, they lead to an reduction of precision. To tackle this challenge, we present an efficient inference approach on the basis of depthwise cluster quantization and model compression. Our method first divides the weights individually across channels and cluster them to generate centroids. Next, we quantize the weights to enforce weight sharing, finally, we retrain the quantized weights in model to adjust quantized centroids value for the purpose of precision enhancement. The experiments on ResNets with CIFAR datasets, indicate that our method reduced the size of ResNet-20 by 4× from 1.2MB to 295KB and 8× from 1.2MB to 152KB with the 5-bit width and 4-bit width, again with no loss of accuracy.