Advanced Quantization Methods for CNNs
Norbert Mitschke, Michael Heizmann · 2020
In this article, methods are presented that improve a quantization procedure for the inference of convolutional neural networks with respect to the resulting accuracy and resource requirements. The existing procedure is extended by pre-training, fine tuning and a suitable compression method. With the presented methods we get better convergence properties for deep quantized networks, we increase the accuracy by a several percentage points and we reduce the number of parameters by up to 90% depending on the model used.