ATA -Net: Asynchronous Temporal Attention for Real-World Multi-Modal 3D Object Detection

Yadong Wang, Sihan Chen, Tianyu Shen, Yongqiang Deng, Juanjuan Li, Kunfeng Wang · 2024

Multi-modal fusion for 3D object detection exploits the benefits of various sensors, mitigating uncertainties arising from single-source sensor, rendering it crucial for visual perception in intelligent transportation. Addressing multi-modal fusion in real-world traffic scenes remains a significant challenge for existing methods, primarily due to issues such as temporal misalignment of data during actual collection and the sparsity problem associated with low-beam LiDAR or roadside LiDAR. To solve the above challenges., this paper proposes an asynchronous temporal attention model (ATA-Net) for achieving more robust multi-modal 3D object detection in real-world traffic scenes. To deal with the spatio-temporal asynchronous multi-modal data, a bi-directional interaction between image and sequential point clouds is realized in ATA-Net, enabling effective global and local feature fusion with spatio-temporal correlations. To mitigate the sparsity concern in point clouds within asynchronous multi-modal data, a object-level densification scheme, with an integration of visual guidance and optical flow estimation, is further designed for producing object-level pseudo point clouds. Experimental validation on the VANJEE real roadside dataset and the V2X-Seq temporal dataset demonstrates the effectiveness of proposed model. Our code will be available at https://github.com/BUCT-IUSRC/Research_ATA-Net.

Read the paper · More papers on PaperTik