Non-Uniform Quantization and Pruning Using Mu-law Companding

SeungKyu Jo, JongIn Bae, Sanghoon Park, Jong-Soo Sohn · 2022

Contemporary deep learning models require high computation costs and large model sizes to realize high accuracy, which is not suitable for limited hardware resources such as mobile or edge devices. Model compression methods such as a quantization that reduces the precision of weights or activation and pruning that removes unimportant nodes have been proposed. However, these methods degrade the accuracy significantly. To overcome this limitation, we propose a non-uniform quantization and pruning method using mu-law companding, which preserves accuracy while simultaneously performing pruning and quantization. Experimental results using the ResNet-18 model on the ILSVRC2012 dataset showed that the accuracy increased by 0.4% in 5-bit and 0.3% in 4-bit, compared to FP32. In addition, we identified areas containing important information by changing the quantization interval, and visually demonstrate why the quantization model outperformed FP32 with Grad-CAM++.

Read the paper · More papers on PaperTik