Target-specific and Temporal Transformer for Visual Tracking

Jiapeng Hu, He Zhao, Jiong Jia, Youming Chen, Liang Zhao, Yamin Han, Meili Wang · 2024

Visual tracking can obtain the location of the interest objects, which is beneficial for 3-D reconstruction in different VR applications. With the progress of deep learning, transformer-based trackers have achieved a remarkable performance gain. However, most existing transformer-based trackers only consider the per-frame localization accuracy, neglecting the potential temporal dependencies among multiple video frames. To address the above issues, we propose a target-specific and temporal transformer for visual tracking. It propagates preceding target-background dynamics into succeeding frames for more coherent tracking results. Firstly, we propose a target-specific candidate generation module to detect target candidates, which enables generated target candidates containing more target-specific information by selecting appropriate search tokens to interact with template tokens. Then, a spatio-temporal correlation transformer is designed to model the temporal evolution of target candidates’ trajectories. It can effectively model the historical temporal target and background information during the tracking for increasing the discriminability and robustness. Extensive experiments have shown that our tracker outperforms previous state-of-the-art trackers on three tracking benchmarks including LaSOT, UAV123, and NFS30.

Read the paper · More papers on PaperTik