Occlusion-Aware Discriminative Networks for Visual Object Tracking
Changhai Wang, Shengjie Zhao, Rongqing Zhang · 2020
Recently, discriminative networks for visual object tracking have been prevailing in visual tracking community due to their powerful discriminative ability. However, when the target is occluded by other semantic or non-semantic backgrounds, most approaches can be fragile since they merely rely on the target appearance to estimate the location of the object in each frame. In this paper, we firstly design a cascaded pyramid module to fuse different parts of target feature map in order to resolve semantic occlusion and improve its robustness. Meanwhile, we propose a spatial attention block which can effectively provide the prior structure information for our module to handle non-semantic occlusion. We then integrate these two proposed module into discriminative networks and training them in an end-to-end fashion. Extensive experiments on OTB-2015, VOT2018, and TrackingNet benchmarks demonstrate that our approach can not only effectively handle occlusion situations but also achieve state-of-the-art performance while running at more than 30 FPS.