THT-Net: A Novel Object Tracking Model Based on Global-Local Transformer Hashing and Tensor Analysis
Shirui Tian, Ningshu Li, Mingxing Duan · IEEE Transactions on Vehicular Technology · 2025
The object point clouds acquired by the original LiDAR are inherently sparse and incomplete, resulting in suboptimal single object tracking (SOT) precision for 3D bounding boxes, especially for small objects at extended ranges. Furthermore, the feature extraction process for object point clouds necessitates substantial computational resources, and the constraints of real-world scenarios often impede the transmission speed essential for practical deployment. Therefore, we propose a novel object-tracking network (THT-Net) based on global and local transformer hashing and tensor analysis, which comprises four key modules. Firstly, point cloud features are extracted using Siamese networks to compute the correlation between the object template region and the search region. Secondly, tensor decomposition and completion are employed to reduce the dimensionality of object point cloud features while preserving and extracting object information. Thirdly, a hash function is constructed to encode the positional information of point clouds, and global and local transformer modules are utilized to analyze the inter-object correlations and predict missing or distant complex objects. Finally, the prediction head module is implemented to reassemble the feature tensor and localize the object-bounding box, thereby finalizing the tracking process. We conduct numerous experiments on the KITTI and nuScenes datasets to evaluate the performance of THT-Net. The experimental results demonstrate that: 1) The THT-Net achieves average success rates and accuracy of 65.5% and 84.1% (KITTI), and 51.36% and 59.88% (nuScenes), respectively; 2) The THT-Net significantly reduces the algorithm training time to 2.06 hours (KITTI) and 3.85 hours (nuScenes). The source code is accessible athttps://github.com/ShiruiTian/THTNet.