SDVRF: Sparse-to-Dense Voxel Region Fusion for Multi-Modal 3D Object Detection

Binglu Ren, Jianqin Yin · 2024

In the perception task of autonomous driving, the performance of multi-modal methods is usually limited by the sparsity of the point cloud or the noise problem caused by the misalignment between LiDAR and the camera. To solve these two problems, we present a new concept, Voxel Region (VR), which is obtained by projecting the sparse local point clouds in each voxel dynamically. And we propose a novel fusion method named Sparse-to-Dense Voxel Region Fusion (SDVRF). Specifically, the sparse point features are fused with dense image features within each Voxel Region. Furthermore, we propose a multi-scale fusion framework to extract more contextual information and capture the features of objects of different sizes. Experiments on the KITTI dataset show that our method improves the performance of different baselines, especially on classes of small size, including Pedestrian and Cyclist.

Read the paper · More papers on PaperTik