Sparse Transformer Visual Tracking Network Based on Second-Order Attention
Xiaolin Yang, Zhiqiang Hou, Fan Guo, Sugang Ma, Wangsheng Yu, Xiaobao Yang · 2024
The Transformer-based visual tracking has demonstrated exceptional performance, but there is still space for further improvement in target feature expression. To address this issue, this paper proposes a Sparse Transformer visual tracking network based on second-order attention to enhance the feature expression capability of the tracking algorithm. Firstly, the spatio-temporal motion information is integrated in this paper to model the motion of the target in the video sequence, thereby enhancing the feature expression of the target. Secondly, the proposed method improves the Transformer structure in the visual tracking by utilizing a sparse self-attention mechanism that focuses on the most crucial information in the target area. More importantly, to further enhance the discriminative ability of the tracker for the target, a mixed second-order pooling (MSOP) module is embedded in the encoder-decoder structure. The proposed method has been extensively evaluated on multiple datasets, achieving success rates of 70.9%, 67.9%, and 64.8% on OTB100, UAV123, and LaSOT, respectively. Furthermore, the method obtains EAO score of 0.485 on VOT2018. The experimental results demonstrate that the proposed method has better tracking performance and is more robust in various complex scenarios.