Post Training Weight Compression with Distribution-based Filter-wise Quantization Step

Shin‐ichi Sasaki, Asuka Maki, Daisuke Miyashita, Jun Deguchi · 2019

Quantization of models with lower bit precision is a promising method to develop lower-power and smaller-area neural network hardware. However, 4- or lower bit quantization usually requires additional retraining with labeled dataset for backpropagation to improve test accuracy. In this paper, we propose a quantization scheme with distribution-based filter-wise quantization step without labeled dataset. ResNet-50 model with 8-bit activation and 3.04-bit weight precision quantized with the proposed techniques achieves top-1 inference accuracy of 74.3% on ImageNet.

Read the paper · More papers on PaperTik