Ensemble Long Short-Term Tracking with ConvNeXt and Transformer
Yuhua Xiao, Yifeng Zhang, Pengyu Ni · 2022 7th International Conference on Image, Vision and Computing (ICIVC) · 2022
Visual object tracking is an important research topic in Computer Vision. The widely used Siamese network architecture learns a similarity metric between target objects and search regions, and locates the targets in video sequences. In this paper, we present an ensemble long short-term tracking algorithm based on ConvNeXt and Transformer. Firstly, a Siamese network with the ConvNeXt backbone is applied to extract features for both target and search regions. Secondly, an encoder-decoder transformer is introduced to capture global feature dependencies. In addition, an IoU-confidence-based tracking ensemble algorithm is designed to capture both long-term stable appearances and short-term variable appearances of the target. The proposed tracker, called STARK-NeXt, achieves a success rate of 68.9% on LaSOT, outperforming STARK by 1.8%.