Instance Fusion for Addressing Imbalanced Camera and Radar Data
Haohui Zhu, Bin‐Jie Hu, Chen Zhao · 2024
Multi-modal fusion is crucial for accurate 3D object detection in autonomous driving, as it directly impacts subsequent planning and decision-making. We fuse camera and millimeter-wave radar for 3D object detection, considering the robustness of radar in adverse weather. However, fusing sparse geometric information from radar with dense semantic information from images is challenging. Previous works extract radar features using the voxelization method following LiDAR, which is often ineffective and may result in information loss due to the sparsity of radar point clouds. To overcome this, we propose an instance fusion module to effectively integrate the two modalities. Leveraging the instance characteristics of queries, we associate and fuse them with radar voxel instances in geometric space, which significantly improves the detection accuracy of objects in velocity measurement, resulting in a 32.9% reduction in velocity error. In addition, radar data is used to enhance the perception awareness of the foreground on images and a cross-attention module is employed to adaptively fuse the features of both modalities. Through the above operations, we fully exploit the multi-level information from radar and validate the effectiveness of fusion at different stages. Our model achieves an NDS of 49.7% and a mAP of 41.5% on the nuScenes dataset.