Joint Language Prompt and Object Tracking
ZhiMin Weng, Jinpu Zhang, Yuehuan Wang · 2024
Recently, utilizing natural language descriptions to assist in object tracking is becoming a trend. However, The tracking performance is compromised with the absence of the Vision-Language (VL) datasets. In this paper, we develop Joint Language Prompt and Object Tracking (JLPT), the first application of prompt learning in VL tracking. JLPT preserves knowledge from large-scale pretrained vision-based models and adopts language descriptions as prompts to aid tracking with a limited amount of datasets. Specifically, the Multimodal Fusion Prompter (MFP) constructs precise language prompts by dimension compression and correlation operation on heterologous modal. It can adaptively reinforce target features and attenuate interference features. Additionally, the Language Rectification Moudle (LRM) is introduced to dynamically adjust language prompts based on target appearance variations, providing temporal adaptability of the language prompts. JLPT achieves notable performance with minimal trainable language prompt parameters and limited VL datasets. Extensive experiments on TNL2K, LaSOT, LaSOText, and OTB99-L confirm the superiority of JLPT in VL tracking over existing trackers.