Post-training 4-bit Quantization of Deep Neural Networks

Sujay S Tadahal, Gopal Bhogar, S. M. Meena, Uday Kulkarni, Sunil V. Gurlahosur, Shashidhara B. Vyakaranal · 2022 3rd International Conference for Emerging Technology (INCET) · 2022

In the past few years, there have been major advancements in the applications of machine learning in various fields. Deep Neural Networks(DNNs) being the most preferred machine learning technique, has evolved with higher implementational complexity and thus require substantial memory bandwidth and storage for intermediate computations, apart from significant computing resources. Network pruning to eliminate trivial connections, optimized DNNs designed for resource-constrained platforms, and reducing the model size of offloading precision requirements for weights and activations are some of the approaches to deploying DNN models on edge devices. DNN quantization has remarkable advantages in bringing machine learning and Artificial Intelligence(AI) from high computing devices to embedded platforms. Mathematical calculations involving integers are much more efficient than Floating-point operations but many existing solutions require a full dataset to regain the accuracy lost during quantization. Many researchers have contributed to the 8-bit quantization scheme that allocates a fixed number of bits for every Channel irrespective of its distribution, and one such implementation is the TFlite optimization toolkit provided by TensorFlow, But the domain of non-uniform quantization with fewer bits than 8 is an active area of research. In this paper, we propose a post-training 4-bit quantization scheme that decides the number of bits allocated to a channel based on its distribution. We also provide methods to calculate the bit requirement for the channel and overcome bias in mean and variance after quantization. Incorporating these two methods our proposed quantization scheme achieves float comparable accuracy. With the proposed quantization, depending on the architecture, there will be around 30-75% reduction in the model size and the accuracy loss will be around 0.9-5.2% as compared to the float model.

Read the paper · More papers on PaperTik