Visual-language modal hybrid tracking algorithm
Long Cheng, Rui Li · 2023
Visual and linguistic modalities provide complementary information for various computer vision tasks. This paper proposes a novel approach for improving the tracking algorithm SiamBAN by integrating visual and language modalities. We introduce the concept of a visual-language modal mixer, which combines visual features and language representations to enhance the tracking performance. Specifically, we leverage a language model to extract semantic features from language descriptions and align them with visual features using a linear layer. The VL Modal Mixer is implemented through the Hadamard product operator, preserving spatial information. The mixed features are then fused with visual features through a residual connection to retain fine-grained visual details. Extensive experiments on benchmark datasets demonstrate the effectiveness of our proposed method, achieving state-of-the-art performance in terms of accuracy and robustness. Our work contributes to the advancement of multimodal tracking algorithms and opens up new possibilities for integrating visual and linguistic cues in computer vision tasks.