ACTrack: Visual Tracking with K-est Attention and LG Convolution
Xuli Luo, Yu Yao, Jianhui Guo, Yuxuan Liu, Liangkang Wang · 2023
In visual object tracking tasks, the long-range dependencies between image features are critical for target localization. Due to its ability to capture the long-range relationships of image features, the attention mechanism has replaced the convolution operation as the dominant way of image representation learning. However, previous Transformer-based trackers use the attention mechanism to capture feature dependencies from the global perspective, inevitably ignoring local feature dependencies and therefore reducing the discriminability of foreground and background, resulting in unreliable tracking in complex sceneries. In our work, we propose a novel and lightweight feature representation approach that combines a global feature representation based on the attention mechanism with a local feature representation based on the convolution operation. Additionally, we employ the K-est attention mechanism to focus on the primary target information and mitigate interference from secondary background information, thus improving tracking performance. Furthermore, lightweight LG convolution enhances the local representation by virtue of inductive bias. Our proposed tracker, named ACTrack, outperforms all state-of-the-art trackers on GOT-10k, LaSOT, TrackingNet, and UAV123 datasets, running at 38 FPS. Notably, without compromising tracking speed, the tracking performance of our ACTrack is significantly superior to that of trackers using only the attention mechanism.