Quantizaiton for Deep Neural Network Training with 8-bit Dynamic Fixed Point

Yasufumi Sakai · 2020

Recent advances in deep neural networks have achieved higher accuracy with more complex models. Nevertheless, they require much longer training time. To reduce the training time, training methods using quantized weight, activation, and gradient have been proposed. Because neural network calculation by integer format improves the energy efficiency of hardware for deep learning models, training methods for deep neural networks with fixed point format have been proposed. However, the narrow representation range of the fixed point degrades the accuracy achieved when using fixed points. In this work, we propose quantization method without accuracy degradation for deep neural networks training. The proposed quantization method can change the fixed point representation range to preserve accuracy by adding bias to the exponent of fixed point. We investigated the effectiveness of proposed quantization method on the ImageNet task using ResNet-50. Using the proposed method, the evaluated model can be trained using 8-bit dynamic fixed point without accuracy degradation.

Read the paper · More papers on PaperTik