Transformer-Based Tracker Fusing Two-Dimensional Semantic Features

Yongsheng Luo, Zhibin Du, Qianling Guo, Zhide Guo · 2025

In recent years, with the rise of Transformer network, the tracking network based on Transformer network architecture has become one of the mainstream frameworks in the field of single target tracking. However, the performance of Transformer-based trackers still encounters certain bottlenecks in complex scenarios such as target appearance changes. Firstly, most methods still adopt independent template learning, lacking communication between the template and the search area; on the other hand, in feature extraction and fusion, the traditional attention mechanism lacks the ability to simultaneously capture global information in both height and width spatial dimensions, which limits the expressive power of features. To address these issues, improvements were made based on the ROMTrack model. By integrating three typical object modeling methods and designing a global grouped coordinate attention module during the feature extraction phase, the new approach combines multi-dimensional global information with attention mechanisms, capturing global in-formation across height and width dimensions, enhancing the comprehensiveness of feature learning and the robustness of object modeling. Experimental results on GOT-10K, LaSOT, LaSOText, TrackingNet, OTB100, and UAV123 datasets show that compared to ROMTrack, there are improvements in both precision and success rate, demonstrating excellent tracking performance.

Read the paper · More papers on PaperTik