CNNBooster: Accelerating CNN Inference with Latency-aware Channel Pruning for GPU

Yuting Zhu, Hongxu Jiang, Runhua Zhang, Yonghua Zhang, Dong Dong · 2022

Channel pruning is one of the mainly used meth-ods in current network model compression. However, existing channel pruning methods lack effective hardware runtime latency guidance, making the reduction in model size not fully converted into inference latency reduction, which results in poor inference performance of the pruned network models. This paper proposes a CNN channel pruning framework, CNNBooster, that incorporates hardware runtime latency in-formation, whose two major contributions are as follows: First, CNNBooster automatically analyses the latency behavior of various CNNs with channel reduction on different GPU hard-ware platforms, which achieves efficient localization of pruna-ble coordinates. Second, CNNBooster uses a flexible grained latency-aware and param-aware pruning algorithm based on prunable coordinates to achieve a significant reduction in model inference latency and parameter size. We evaluate CNNBooster with three benchmark models (VGG16, ResNet18 and MobilenetV1) on three NVIDIA-series platforms (V100, RTX 2080 Ti and Jetson Nano). The experimental results show that CNNBooster can achieve a maximum 67.51 % latency re-duction and 1.4x performance improvement compare with current state-of-the-art works.

Read the paper · More papers on PaperTik