Accelerating CNN Inference: Hybrid Model Pruning for Reduced Computational Demands and High Accuracy

Abdulgader Ramadan Gadah, Suhaila Isaak, Abdul‐Malik H. Y. Saad · 2024

Convolutional Neural Networks (CNNs) have revolutionized various fields, including image recognition, autonomous vehicles, and medical imaging. These models are computationally expensive by including irrelevant processing if the model also affects the system’s accuracy. Model pruning is one of the practical approaches to deal with the computation of the model. This paper proposed the hybrid pruning technique using the EfficientNetB3 model, which achieved a remarkable 50% reduction in model parameters and a 10% decrease in inference time while maintaining an accuracy of over 96% (96.09% vs. 96.22% for the unpruned model). The hybrid approach also demonstrated a significantly faster execution time of 2864.18 seconds, outperforming single pruning methods (2982.77 seconds) and the unpruned model (3197.25 seconds)—this substantial reduction in computation time positions hybrid pruning as the most effective strategy for enhancing CNN performance. The proposed approach provides effective results while reducing the overall computation of the model.

Read the paper · More papers on PaperTik