Visual Tracking with Temporal Contextual Attention

Qin Li, Wenjie Zou · 2023

Due to changes in the appearance, scale, and other aspects of tracking objects, accurate tracking has great difficulties in the current existing visual trackers. Recently, there has been significant progress in time-domain trackers, which has inspired us to solve this problem. Therefore, in this paper, we effectively address the issue of object scale and appearance changes during the tracking process from the perspective of mining temporal context information. This paper proposes Visual Tracking with Temporal Contextual Attention (TCFormer). The modeling scheme not only extracts the target-specific distinguishing features and the association between the target and the search area, but also effectively integrate temporal and spatial context information. Specifically, the proposed method includes a temporal context feature extraction module (TCE) and a extraction and integration feature module (EIF). TCE uses attention mechanism to effectively fuse temporal and spatial context information of keyframes. EIF is utilized to extract features from templates and search regions in multi-stage feature networks, as well as effectively combine the template and search regions features. By multi-level matching of templates and search features, non-target features can be suppressed to achieve discriminative feature extraction. Experimental results on visual tracking benchmarks including OTB, UAV123, GOT-10K, LASOT and TrackingNet demonstrate that TCFormer achieves the convincing performance outperforming most state-of-the-art trackers.

Read the paper · More papers on PaperTik