HA-Tracker: A Hybrid Architecture Tracker with Spatiotemporal Mamba Motion Model for UAV-Based Video Multi-Object Tracking
Pengfei Zhang, Leigang Sun, Chang Li, Qinyi Wang, Qingtao Hao, Junjing Lu, Lu Zuo, Xiaoqian Ma · Remote Sensing · 2025
UAV-based video multi-object tracking (MOT) is a significant task in the field of remote sensing. However, current research still faces critical issues: (1) the limitations of the single architecture of DNNs inherently hinder performance improvement of object detection; and (2) current linear modeling approaches for spatiotemporal relationships fail to capture complex motion patterns in the real world. To overcome the aforementioned issues, a hybrid architecture tracker (HA-Tracker) with a spatiotemporal Mamba motion model for UAV-based video MOT is the first to be proposed, which has the following innovations and contributions: (1) a CNN–Transformer–Mamba detector (CTM detector) is proposed to enhance the capability of object detection, which is a novel synergistic fusion framework for simultaneously fusing the local details of a CNN, the global context of a Transformer, and the long-range dependency of Mamba; and (2) a spatiotemporal Mamba motion model (STM3) is proposed to improve tracking accuracy by modeling the nonlinear spatiotemporal motion relationships of object trajectories. Extensive experimental results indicate that our HA-Tracker achieved outstanding performance, with multiple object tracking accuracy (MOTA) metrics of 44.76% and 52.22% and identity F1 scores (IDF1) of 60.33% and 72.34% on the Visdrone and UAVDT datasets, respectively. These results validate the effectiveness of HA-Tracker, which outperforms the existing MOT networks.