Video Object Detection Using Motion Context and Feature Aggregation

Jaekyum Kim, Junho Koh, Jun Won Choi · 2020

The deep learning technique has recently led to significant improvement in object-detection accuracy. Numerous object detection schemes have been designed to process each frame independently. However, in many applications, object detection is performed using video data, which consists of a sequence of image frames. Thus, the object detection accuracy can be improved by exploiting the temporal context of the video sequence. In this paper, we propose a novel video object detection method that exploits both the motion context of the object and spatio-temporal aggregated features to enhance the video object detection performance. First, the motion context of the object is extracted by the correlation operator between the feature maps of two adjacent frames. In addition to generating the motion context, the spatial feature maps for N adjacent frames are aggregated to boost the quality of the feature map with gated attention network.

Read the paper · More papers on PaperTik