Benchmarking of Quantization Libraries in Popular Frameworks
Tejas Dubhir, Mayank Mishra, Rekha Singhal · 2021 IEEE International Conference on Big Data (Big Data) · 2021
Quantization is a technique to reduce the size and computation time of machine learning models by reducing the precision of model parameters. However, quantization may reduce the accuracy of the model—popular ML frameworks such as Pytorch and Tensorflow support the quantization process. A quantization process may differ in reducing the number of bits (8, 16, or 32) of model parameters, static, dynamic, and quantization aware training. In this paper, we evaluate the various features of the quantization process supported in Pytorch and Tensorflow on CNN and GNN based Recommendation models. We have compared Pytorch and Tensorflow quantization libraries for the memory efficiency and accuracy of the quantized models. The paper also presents and discusses an additional challenge of quantization in GNN based Recommendation models having embedding layers.