Calibration Data-based CNN Filter Pruning for Efficient Layer Fusion
Krishna Teja Chitty-Venkata, Arun K. Somani · 2020
Convolutional Neural Network (CNN) optimization is critical to reduce the inference latency on computing devices like CPU and GPU. The most important step in efficiently executing these algorithms involves combining multiple operations within a single convolutional layer through a process called Layer Fusion. We first form a correlation between Layer Fusion and data management on computing platforms like CPUs and GPUs, and analyze its significance for different networks under different model compression techniques, e.g., Pruning and Quantization. Weight Pruning removes redundant parameters in the network thereby shrinking the model size. This method, however, creates sparse matrices which can hamper the performance on CPU and GPU devices. Several node/filter pruning algorithms have been developed to resolve the bottlenecks of irregular pruning and reduce the inference time. Although symmetric pruning techniques reduce the forward path computation time, the weights and activations are still executed in floating point precision which can be further optimized by Quantization. This method reduces the bit-width of individual CNN parameter from high precision (Float32) to lower precision (Int8). Even though several model compression (pruning and quantization) techniques have been developed to improve performance, their integrated study with Layer Fusion has not been performed. In today's scenario, not all CPUs and GPUs explicitly support Int8 multiplication. Hence, we analyze the performance of optimized implementation of Int8 Quantization on such devices like Intel's Skylake and Nvidia's Tesla V100. We develop a novel Node Pruning algorithm to remove redundant filters which can aid in efficient implementation of Layer Fusion/Quantized Networks. We compare the execution time of our combined pruning and quantization implementation with a traditional node pruning algorithm which achieved a mean speedup of 3.5 times on Skylake CPU.