Exploiting Multimodal Prompt Learning and Distillation for RGB-T Tracking
Qingkuo Hu, Yichen Li, Wenbin Yu · 2025
RGB-Thermal (RGB-T) multimodal tracking has gained widespread attention due to its robustness in handling complex scenarios. Some existing methods focused on fully fine-tuning RGB-based trackers, which was parameter-inefficient and prone to overfitting due to the scarcity of multimodal data. Therefore, recent studies have explored multimodal prompting strategies, which mainly treat RGB as the dominant modality and TIR as an auxiliary prompting modality. This asymmetric framework lacks adaptability in changing dominant modalities, resulting in reduced robustness. To address these limitations, we propose LRPD: a novel multimodal Low-Rank (LoRA) Prompting and Distillation tracking framework. Specifically, the framework consists of two distinct stages. In the first stage, we pre-train a teacher (ViT-L encoder) and a student model (ViT-B encoder) using our LoRA-Prompting (LoRA-P) module. LoRA-P adopts a symmetric architecture that enables bidirectional cross-modal interaction in a parameter-efficient manner. In the subsequent stage, we design a prompt-driven knowledge distillation framework to transfer knowledge from the large teacher model to the lightweight student model. Task-specific designs, including an enhanced patch masking strategy, a feature alignment projector, deep visual prompts, and specialized distillation losses, are tailored to optimize student's tracking performance. Extensive experiments on three popular RGB-T tracking benchmarks demonstrate our method achieves new state-of-the-art performances.