CTTrack: Leveraging Cross-Transformer for Multi-Scale Feature Fusion in Multi-Object Tracking

Ziyang Zhang, Chuqing Cao, Fangjun Zheng · 2024

Multi-object tracking is a fundamental and challenging task in the field of computer vision. Contemporary multi-object tracking methods mainly follow the detection-tracking paradigm, in which image features play a crucial role in tracking. However, most existing object tracking methods use displacement regression or motion estimation to determine object positions, which relegates the importance of features to a secondary role. Moreover, single feature information is insufficient to effectively distinguish similar objects. Therefore, this paper proposes CT-Track, a target tracker that integrates semantic and structural information. Specifically, we densely sample candidate regions of hundreds of objects in image pairs for contrastive learning and introduce a cross-transformer to fuse rich structural information from low-level features with semantic information from high-level features. This fusion enhances localization capabilities and the understanding of complex scenes, resulting in more robust tracking performance. Experiments on the MOT17 and joint datasets highlight the benefits of multi-scale feature fusion using cross-transformer in multi-object tracking.

Read the paper · More papers on PaperTik