TandemFuse: An Intra- and Inter-Modal Fusion Strategy for RGB-T Tracking
Xinyang Zhou, Hui Li · 2024
Visual object tracking is a prominent task in the field of computer vision, with significant potential in autonomous driving, human-computer interaction, and intelligent surveillance. Many studies have focused on tracking using single-modality data. Among these, the RGB modality is renowned for its robust color and detail capture capabilities, yet it is susceptible to motion blur, occlusions, and low-light conditions. Conversely, the TIR modality can overcome these issues but is limited by lower resolution and higher costs, making target recognition in complex scenes more challenging. With the rise of multi-modal learning, integrating RGB and TIR modalities can significantly enhance the robustness of single-modality trackers across various scenarios. This paper proposes a comprehensive multi-modal learning strategy for fusing RGB and TIR modality in the tracking task. Two key modules are designed: intra-modal data fusion and inter-modal data fusion. For intra-modal data fusion, we utilize feature pyramid techniques to merge multi-scale representations from single modalities. Subsequently, features learned from both modalities are sent to inter-modal data fusion for enhanced tracking. Our model is trained on the RGBT234 dataset and tested on the GTOT dataset, achieving a success rate of 0.454 and a precision rate of 0.438. The insights and methodologies derive from this research offer guidance for future multi-modal object tracking studies and underscore the critical role of intra-modal fusion in enhancing the efficiency of multi-modal integration.