SAFF: Multi-Modal 3D Object Detection with Search Alignment and Full-Channel Fusion
Shuqin Zhang, Yadong Wang, Yongqiang Deng, Juanjuan Li, Kunfeng Wang · 2023
Multi-modal object detection is a core component of grounded applications of autonomous driving and intelligent transportation. Most of the existing multi-modal object detection methods rely on the calibration of datasets to different modal-ities to find the correspondence between modalities for fusion. However, in realistic scenarios, there is inevitably the problem of temporal and spatial asynchrony of datasets, which hinders the fusion of multi-modal features. In addition, since features are often augmented and aggregated, it is difficult to effectively combine the advantages of both modalities using only the original calibration parameters. In this paper, we propose the SAFF model to improve the multi-modal object detection algorithm from two aspects, i.e., feature alignment and feature fusion. For feature alignment, the multi-scale local search module finds the optimal image features for point cloud matching by iterative learning to capture more important image semantic information. To effectively fuse multi-modal features, the feature fusion strategy is optimized by introducing information of different dimensions of the features. Compared with the baseline, our method achieves 3.38% higher performance in 3D detection on the roadside VANJEE dataset with spatio-temporal asynchrony problems. In addition, to verify the effectiveness of the method with full-channel fusion, experiments are conducted on the public KITTI dataset and our method achieves 2.23% and 2.26% higher performance for pedestrians and cyclists respectively.