Bi-directional Enhanced Network of Cross-Modal Features for RGB-T Tracking

Hang Zheng, Yuanjing Zhu, Peng Hu, Honglin Wang, Hongwei Ding · 2024

The current deep learning based RGB-T target tracking algorithm first uses convolutional neural networks to do feature extraction on the sample frame and then classifies the target and the background, which achieves better tracking results on generalized datasets, but still has some limitations in dealing with the challenge of background clutter (BC). Background clutter refers to the presence of objects around the tracked target that interfere with the tracking, making it impossible for the tracker to accurately distinguish between the target and other interfering objects. To address this problem, we propose a cross-modal feature bi-directional enhancement network, which first derives the correlation between two modal data pairs through the channel attention mechanism to realize the sharing of multi-modal information, and then enhances the position discriminative features of the target in the other modality by using the position discriminative features of one modality through the spatial attention mechanism to improve the accuracy of localization. Experiments on two RGB-T benchmark datasets verify the effectiveness of the algorithm.

Read the paper · More papers on PaperTik