Video Object Detection Based on Deformable Convolution
Yanjun Liang, Zhiming Wang · 2020
Video object detection is one of the important research directions in the field of computer vision with applications in various domains, e.g. public security, traffic management, etc. Nevertheless, it is very challenging to extend the image-based object detection method to video object detection. In the existing methods, the feature information of multiple adjacent frames is aggregated into the current frame to improve the detection result. However, these methods only sample the information of adjacent frames on the sampling points obtained by prior knowledge or by calculating explicit motion references. Therefore, the feature information obtained by this sampling method is often insufficient. Furthermore, if the selected sampling points are inaccurate, the original feature information may be lost. In this paper, we use modulation-based deformable convolution for feature sampling, and the sampling positions and weight are generated through additional convolutional layers. This not only enriches the sampling information but also minimizes the intervention of prior knowledge. According to experiments on the Image VID dataset, the proposed method significantly improve detection accuracy compared with existing methods.