Pseudo-depth Guided Multi-object Tracking
Zijie Zhuang, Xiaokai Yi, Shuaixiong Hui, Jing Yang, Hanli Wang · 2025
Multi-object tracking (MOT) constitutes a critical task in video analysis, requiring precise detection of targets across consecutive frames and robust association of their identities over time. Existing methods predominantly rely on extracting visual cues from the single RGB modality for target tracking, which demonstrates satisfactory performance in controlled environments. However, such unimodal approaches exhibit significant limitations in complex scenarios, particularly under challenging conditions such as severe occlusions, poor illumination, and adverse weather conditions, where the reliability of RGB data is substantially compromised. In this paper, a pseudo-depth guided multi-object tracking (PDMOT) method is proposed that effectively addresses the limitations of conventional RGBbased tracking methods through innovative depth feature fusion. The proposed method introduces a pseudo-depth generation mechanism that harnesses the generative capabilities of pretrained large-scale models, thereby enabling robust feature representation in complex scenarios. By strategically integrating both RGB modality and depth information, the PDMOT framework achieves enhanced tracking performance while maintaining computational efficiency. The experimental results indicate that the proposed PDMOT method yields substantial improvements, outperforming existing state-of-the-art (SOTA) approaches by 2.3% in HOTA and 2.7% in AssA on the DanceTrack dataset. These results underscore the effectiveness of PDMOT in enhancing both tracking accuracy and robustness. Furthermore, its superior association performance on the MOT17 dataset further corroborates that the proposed method outperforms other Transformer-based techniques.