Real-Time Object Detection by Feature Map Forecast for Live Streaming Video
Masato Fujitake, Akihiro Sugimoto · 2021
This paper proposes a method that jointly learns to detect objects at the current frame and forecast the next frame’s future feature map. Previous offline detectors have shown the effectiveness of utilizing future information in video object detection; however, we cannot take such an approach when dealing with live streaming videos. In contrast, we utilize the forecast feature map with the current and past frame feature maps for object detection, where forecast feature maps are learned using observation of the present and past frames. To maintain a reliable forecast, we introduce a scheduler network, which decides whether we use the forecast feature map as input or extract the feature map from the next frame. Evaluations of our proposed model on the ImageNet VID dataset demonstrate the superior performance of our model against the public benchmark at similar architectures, with achieving 65.7% mAP at 38.9 fps.