Reducing Memory Requirements of Convolutional Neural Networks for Inference at the Edge

Tomáš Bravenec, Tomáš Frýza · 2021

The main focus of this paper is to use post training quantization to analyse the influence of using lower precision data types in neural networks, while avoiding the process of retraining the networks in question. The main idea is to enable usage of high accuracy neural networks in devices other than high performance servers or super computers and bring the neural network compute closer to the device collecting the data. There are two main issues with using neural networks on edge devices, the memory constraint and the computational performance. Both of these issues could be diminished if the usage of lower precision data types does not considerably reduce the accuracy of the networks in question.

Read the paper · More papers on PaperTik