Demystifying Compression Techniques in CNNs: CPU, GPU and FPGA cross-platform analysis

Remya Ramakrishnan, Aditya Dev, Darshik A.S., Renuka M. Chinchwadkar, Madhura Purnaprajna · 2021

Convolutional Neural Networks (CNNs) are known for their high-performance despite its huge memory requirement and computational complexity. A wide range of compression techniques to reduce the number of parameters and hence computational and memory complexity have been exploring recently. In this paper, we analyse three widely used categories of techniques viz. quantization, pruning and tensor decomposition to make a cross-platform performance comparison on CPU, GPU and FPGA. These techniques are not mutually exclusive and hence can be combined to get better compression and a better speed-up on devices. Our focus is to highlight the contrasting impact of optimization techniques on devices and performance objectives. We observe a speed-up of 3.8 to 15.6× on CPU, 3.4 to 7.2× on GPU and 10.5 to 29.4×on FPGA across the models and compression techniques under consideration. We also achieved a compression of 93 to 97% across models with acceptable accuracy. Blended techniques have shown a better speed-up on FPGA compared to CPU and GPU as the caching effects, memory accesses and compiler optimizations slow down the inference on these general-purpose machines.

Read the paper · More papers on PaperTik