Towards Accurate Low Bit DNNs with Filter-wise Quantization
Hoseung Kim, Lee Kwang-Bae, Dongkun Shin · 2020
Execution of deep neural networks (DNNs) on a resource constraint device has become a rising issue of recent neural network research. Quantization using low bit-widths for networks is one of the most effective compression techniques. Although there have been many studies that use different bit-widths per layer to compress further the model instead of using a single bit-width, they achieved a limited reduction on the parameter size due to the layer-wise bit-width assignment. In this paper, we propose a more fine-grained and multi-precision quantization technique, called filter-wise quantization. Regularization is used while training networks to partition filters into various precision. In experiments, we show that our technique can provide better accuracy at a smaller parameter size at various DNN models for CIFAR-10 and CIFAR-100 data sets.