Performance Evaluation of Stochastic Quantization Methods for Compressing the Deep Neural Network Model

June-Yeong Choi, Joonhyuk Yoo · Journal of Institute of Control Robotics and Systems · 2019

Discretizing the Deep Neural Network (DNN) is enforced by a rounding function of weights and activations, which can severely degrade the performance of DNN. This paper categorizes the available quantization methods to be uniform or non-uniform according to the distribution of its quantized interval, and proposes a noble stochastic rounding method to reduce the accuracy loss of the quantized DNN model. Two stochastic rounding methods of linear or nonlinear probability distribution are quantitatively evaluated and analyzed to give an important insight to employ them for designing the compressed DNN model, by comparing them with the previous deterministic nearest rounding. Experimental results show that the stochastic rounding methods are not always superior to the deterministic one and the optimal performance is obtained via a hybrid rounding scheme when applying the stochastic rounding to the input adjacent to the intermediate value between the neighboring quantized values and still using the deterministic one near to the quantized value.

Read the paper · More papers on PaperTik