TVTracker: Target-Adaptive Text-Guided Visual Fusion for Multimodal RGB-T Tracking
Fang Gao, Wenjie Wu, Yan Jin, Jingfeng Tang, Hanbo Zheng, Shengheng Ma, Jun Yu · IEEE Internet of Things Journal · 2025
Current multi-modal sensor trackers mainly rely on visual cues for target tracking. However, in challenging scenarios, the visual information acquired by multi-modal sensors has limited descriptive ability when the target state undergoes significant changes, which may lead to unsatisfactory tracker performance. In this work, we propose TVTracker, a two-stage tracking framework. It leverages semantic information of the target state to enhance visual cues, enabling effective RGB-T tracking. The first stage is the generation of target text descriptions. Utilizing the Bootstrapping Language-Image Pre-training (BLIP) model, we generate target textual descriptions that match the images in the dataset. In the second stage, text-guided visual fusion is performed for target tracking. Target textual and visual features are extracted separately using the text and visual branches. Then the target textual features are integrated with the visual features to localize the target position and predict the target bounding box. In the text branch, we design the Target Text Adaptation Enhancement (TTAE) module to mitigate the interference of low-quality target textual features on visual features. In the visual branch, we develop the multi-modal visual information prompters, which include the Multi-modal Visual Shared Information Prompter (MVSIP) and the Multi-modal Visual Shared and Complementary Information Prompter (MVSCIP), to facilitate learning of multi-modal shared or complementary visual prompts. Experiments on the LasHeR, RGBT210, RGBT234, and VTUAV datasets demonstrate the effectiveness of TVTracker.