Visual multi-object tracking by interaction network
Richard Irampaye, Junyu Chen, Xiaolin Zhu, Yan Zhou · 2024
Currently, one-shot models are receiving much attention in the field of multi-object tracking (MOT), which combines detection and re-identification (ReID) task in a way that strikes a good balance between accuracy and speed. However, there is a learning difference between detection and ReID task, which tends to undermine the differential learning between them on branching tasks, thus affecting the performance of multi-object tracking. Inspired by this, we propose to solve this problem by employing an interaction network to alleviate the competitive learning between them, so that each branch can learn its task representation. In addition, the different scale features tend to lead to semantic fusion and scale inconsistency problems. We introduce a multi-scale channel attention module to prevent semantic-level misalignment, and enhance single-scale resolution features to improve the embedding capability of features. By combining the above two network modules into a one-shot MOT framework, we construct an online MOT tracker to enhance tracking performance. We conducted benchmark experiments on the MOT16 and MOT20 datasets and achieved advanced performance, demonstrating the effectiveness of the proposed tracker.