CLRFI: A multimodal tracking via contrast learning and region feature interaction

Weidai Xia, Jikun Dong, Xingliang Mao, Fangfang Li · Journal of King Saud University - Computer and Information Sciences · 2026

RGB-Thermal (RGB-T) target tracking algorithms aim to achieve robust all-weather tracking by leveraging the complementary properties of multiple modalities. Previous methods often fuse RGB and TIR search area features directly, which can introduce redundant background noise and lead to inaccurate tracking frame predictions when target features are missing due to occlusion or deformation. Additionally, many methods only fuse depth features, neglecting the potential value of multi-layer shared cues from different modalities. This limitation restricts cross-modal interactions within local regions, resulting in inadequate context modeling and reduced tracking robustness. To address these issues, this study proposes a new RGB-T target tracking algorithm, CLRFI. CLRFI utilizes the Vision Transformer (ViT) architecture for RGB-T tracking scenarios and implements a comparative learning strategy focused on the target region. This approach enables the backbone network to thoroughly explore the unique advantages of targets and their inter-modal relationships across different modalities, thereby optimizing the target’s representation. Furthermore, CLRFI leverages potential correlations among modalities to construct cross-attention, effectively enhancing the interaction of salient features in the search region and improving tracking performance. On three publicly available RGB-T tracking benchmark test sets, CLRFI demonstrates superior performance, particularly in handling bounding box blurring and tracking offsets. Extensive experiments on three public benchmark datasets demonstrate the effectiveness of the proposed CLRFI, achieving Precision and Success Rates (PR/SR) of 83.8%/62.2% on RGBT234, 82.5%/60.9% on RGBT210, and 68.2%/54.2% on LasHeR.

Read the paper · More papers on PaperTik