A Convolutional Neural Network Pruning Method Based On Attention Mechanism
Xiao Jie Wang, Wenbin Yao, Huiyuan Fu · Proceedings/Proceedings of the ... International Conference on Software Engineering and Knowledge Engineering · 2019
Pruning effectively reduces the size of neural networks, which facilitates deployment of neural networks in production environment, especially in embedded systems with limited computing resources.In this paper, we propose a convolutional neural network pruning method based on attention mechanism.We add a attention module to model to generate scaling factors for channels.The scaling factors are considered as channels' importance score, thus filters and convolution kernels corresponding to channels with lower importance score are removed.Our method has the ability to learn importance of channels during training, instead of considering only the direct impact of parameters like existing methods.Moreover, it does not depend on any dedicated libraries, so could be combined with other compression methods for better performance.In experiments, we prune about 90% parameters in VGGNet with 0.67% accuracy drop and prune about 50% parameters in ResNet-56 with 1.02% accuracy drop.