DFNet: A Dual LiDAR–Camera Fusion 3-D Object Detection Network Under Feature Degradation Condition

Tao Ye, Ruohan Liu, Chengzu Min, Yuliang Li, Xiaosong Li · IEEE Sensors Journal · 2025

LiDAR-camera fusion is widely used in 3-D perception tasks. In LiDAR and camera sensing tasks, the hierarchical feature abstraction capability possessed by the deep network is beneficial to capture the detailed information from point clouds and RGB images. However, it tends to filter some of the information to extract important features, where the problem of feature degradation due to loss of useful information is inevitable. The deterioration of LiDAR-camera fusion due to feature degradation, brought about by this factor, becomes a challenging problem. It reduces object recognition and leads to decreased detection accuracy. To address this problem, we propose a dual LiDAR-camera fusion network (DFNet) based on cross-modal compensation and feature enhancement. We design a multimodal feature extraction (MFE) module to complement the sparse features of the point cloud utilizing image features and focusing on the spatial information of the features. Then, we introduce a multiscale feature aggregation (MFA) module to generate bird’s-eye view (BEV) representations of the features, which generates feature proposals that are then input to the voxel-grid aggregation (VGA) module to obtain the grid-pooled features. Meanwhile, the VGA module receives the feature proposals extracted from the image backbone and projects the point cloud through voxels to obtain voxel-fused features. Finally, we aggregate the grid-pooled features and voxel-fused features to produce more informative fused features. The results on the KITTI dataset illustrate that DFNet outperforms most 3-D object detection methods, achieving the 3-D detection performance of 77.88% mAP, which indicates that our method is effective in dealing with feature degradation.

Read the paper · More papers on PaperTik