Visual Object Tracking via CLIP Based Visual Prompt Tuning
Futian Wang, Jie Fang, Siyuan Yang · 2025
Pre-trained large models have recently shown exceptional performance in the field of computer vision, attributed to their complex architectures, extensive parameters, and large-scale dataset pre-training. These models exhibit strong feature representation capabilities, requiring only task-specific fine-tuning for various downstream tasks. This paper introduces the pre-trained CLIP visual backbone into object tracking to enhance accuracy and robustness. However, traditional full-parameter fine-tuning poses a significant computational burden. To mitigate this, we introduce a visual prompt tuning technique that adjusts only a small fraction of parameters-specifically, the prompt vectors, feature interaction module, and prediction head-while keeping the CLIP backbone frozen. Our visual object tracking algorithm modifies just 3.34% of the learnable parameters, greatly reducing computational resource requirements. We also analyzed current feature interaction methods and designed an efficient feature interaction module that promotes effective interaction between template and search region features, enabling precise target localization. Experimental results demonstrate that, compared to current leading tracking networks, our approach not only significantly improves tracking accuracy but also considerably reduces training time and resource consumption. These findings validate the effectiveness of our method and highlight the significant potential of visual prompt tuning in object tracking tasks.