Residue Number Systems Quantization for Deep Learning Inference

Sergey Sivkov · WSEAS Transactions on Computers archive · 2023

Quantization of learned CNN weights to Residue Number System can improve inference latency by taking advantage of fast and precise low bit integer arithmetic. In this paper we review the mathematical aspects of RNS operations for signed integer values and evaluate implementation choices for conversion of conventional float-point PyTorch weights of CNN models to RNS representation. We also present a workflow to convert weights of PyTorch neural network layers specific for computer vision domain to 4-bit RNS moduli-sets able to maintain classification accuracy within 5% of 8-bit quantization baseline.

Read the paper · More papers on PaperTik