“Separate-coupled” hierarchical framework for accurate visual object tracking
Long Liu, Kang Liu, Zhen Wei, Jiaqi Wang · Image and Vision Computing · 2026
In visual tracking tasks, complex and frequent scale changes have a significant impact on accurately estimating the object bounding box. Mainstream full-transformer trackers lack the exploitation of effective information between different scales, limiting their applicability on tracking the object with scale changes and resulting in reduced tracking accuracy. To alleviate the above problem, we propose a “separate-coupled” hierarchical visual tracking framework, named SCHTrack. In detail, we construct a hierarchical feature space based on scale visual transformer and design a “separation-coupled” interaction mode (SCT). It is based on the extraction of shallow-scale features, which are then fed to high-level space for performing the extraction and embedding of the semantic information. To mine more object-aware spatial information, we introduce a spatial scale modulation (SSM) module to enhance the scale attention feature by integrating the information across layers. In addition, a hierarchical feature multiplexing decoder (HMD) is further proposed, which multiplexes hierarchical scale features to decode the shape envelope of the object for obtaining more accurate object bounding box estimation. Extensive experiments on five tracking benchmark datasets demonstrate that our tracker achieves comparable performance. In particular, the proposed method surpasses other advanced trackers and has better scale adaptive ability in the aspect ratio change, deformation, low resolution, scale variation, motion blur, and partial occlusion on the LaSOT dataset.