FFTransMOT: Feature-Fused Transformer for Enhanced Multi-Object Tracking
Xufeng Hu, Younghoon Jeon, Jeonghwan Gwak · IEEE Access · 2023
Multi-object tracking (MOT) is a key task in the field of computer vision, involving the identification, tracking, and classification of multiple objects in videos, connecting their trajectories to form a complete sequence. MOT consists of two fundamental components: object detection and data association, which require detecting objects in each frame, determining the objects to be tracked, performing data association in the next frame, and predicting the future trajectory of the objects. In this paper, we propose a Feature-Fused Transformer for Enhanced Multi-Object Tracking (FFMOT), where the Feature Fusion module is introduced to aid the model in determining object trajectories between frames and predicting their positions in future frames. We use a model to fuse the features extracted by the encoder at framet-1with the featurest, resulting in new features, and the decoder performs data association matching between frametand the fused features. We also employ the self-attention mechanism to capture dependencies between input features, enhancing the accuracy and stability of object detection. We evaluated the proposed model rigorously on four datasets (MOT16, MOT17, DanceTrack, BDD 100k) and the experimental results demonstrate that our proposed FFMOT model outperforms other trackers in terms of tracking accuracy and robustness in MOT.