Visualization and Pruning of SSD with the base network VGG16
Xuemei Xie, Xiao Han, Quan Mi Liao, Guangming Shi · 2017
This paper considers the state of art real-time detection network single-shot multi-box detector (SSD) for multi-targets detection. It is built on top of a base network VGG16 that ends with some convolution layers. Its base network VGG16, designed for 1000 categories in Imagenet dataset, is obviously over-parametered, when used for 21 categories classification in VOC dataset. In this paper, we visualize the base network VGG16 in SSD network by deconvolution method. We analyze the discriminative feature learned by last layer conv5_3 of VGG16 network due to its semantic property. Redundancy intra-channel can be seen in the form of deconvolution image. Accordingly, we propose a pruning method to obtain a compressed network with high accuracy. Experiments illustrate the efficiency of our method by comparing different fine-tune methods. A reduced SSD network is obtained with even higher mAP than the original one by 2 percent. When only 4% of the original kernels in conv5_3 is remained, mAP is still as high as that of the original network.