Exploring Quantization-Aware Training on a Convolution Neural Network
Omar Kaziha, Talal Bonny · 2020
In this paper, a convolutional neural network is trained and tested on the MNIST handwritten digit recognition dataset. Quantization-aware training is performed on the model during training and the effects on model size and accuracy are explored and compared to the full precision model. The model is tested in full precision (32-bit), 8-bit precision, and selected quantized layers (mixed) as well. The results have shown a very slight reduction in accuracy of no more than 0.12%, with a compressed model size by 4X for precision of 8-bits compared to 32-bits model. Quantization of selected dense layers yielded a higher accuracy and a smaller model size than both the 32-bit model and the model with quantization of selected convolution layers, resulting in an MNIST handwritten digits classifier with 99.42% accuracy and a size of 0.2133 Mb.