Learning spatial self‐attention information for visual tracking
Shengwu Li, Xuande Zhang, Jing Xiong, Chenjing Ning, Mingke Zhang · IET Image Processing · 2021
Abstract Visual object tracking has been a fundamental topic in computer vision and many convolutional neural network (CNN) based trackers proposed in recent years have achieved state‐of‐the‐art performance, multi‐domain convolutional neural network (MDNet) is one of the most representative CNN‐based algorithms with high performance. In order to significantly improve the robustness and accuracy of the MDNet algorithm, a multi‐domain convolutional neural network based on spatial self‐attention information learning (SAMDNet) is proposed in this study. The authors use the spatial self‐attention module for the spatial information learning of the model. The spatial self‐attention module selectively aggregates the feature at each position by a weighted sum of the features at all positions. Under the control of self‐learning parameters in this module, spatial attention information can be flexibly acquired. The authors also propose a novel interval loss term to solve the problem of different classes with the same semantics in the training data. Finally, an anomaly detection module is carefully designed for relocation after the algorithm completely lost the target. Extensive experiments on the object tracking benchmark (OTB) and the visual object tracking challenge (VOT) benchmarks show that the proposed tracker outperforming most of the state‐of‐art trackers.