Spatiotemporal Visual Tracker Based on Deformable Attention
Zhuo Jia Wu, Askar Hamdulla, Yucheng Huang · 2024
At present, the visual tracking method based on Transformer has made significant progress, and has shown excellent performance on various data sets. However, the traditional Transformer tracker pays too much attention to global features, and its intensive attention calculation will increase the space-time complexity of the tracking algorithm. And the features of the relevant region will be affected by the features of the non-relevant region, resulting in a decrease in tracking accuracy. In addition, the Transformer tracker lacks the ability to utilize the spatiotemporal information between frames, which further hinders the improvement of its tracking performance. To solve these problems, this paper proposes a new tracker SVT-DA based on deformable attention. The tracker captures local and global important features through CNN and deformable attention. In addition, it uses a cross-attention module nested with self-attention to achieve deep semantic information interaction between template features and search features. In order to verify the effectiveness of the proposed method, we conduct experiments on LaSOT, GOT-10K, OTB100 and VAU123 datasets. The result indicators are 4%~5% higher than the baseline method. The experimental results show that our method improves the tracking performance while reducing the number of model parameters, and is highly competitive with other state-of-the-art methods.