Compression of convolutional neural networks: A short survey
Ratko Pilipović, Patricio Bulić, Vladimir Risojević · 2018
Nowadays, convolutional neural networks (CNN) are considered as the state-of-the-art algorithms for various tasks, especially for image classification and recognition. Because of that, more and more attention is aimed towards implementation of CNNs on embedded systems. Main obstacles for implementing CNNs on embedded systems are their large model size and large number of operations needed for inference. In order to surpass these obstacles, algorithms for CNN compression tend to lower model size and number of operations needed for inference. In this paper we review the state-of-the-art in CNN compression. To this end we divided all approaches for CNN compression into three groups: precision reduction, network pruning and design of compact network architectures. After presenting the main approaches in each group we conclude that the future CNN compression algorithms should be co-designed with hardware which will process deep learning algorithms.