MRMCFusion: Robust 3D Object Detection via Multi-represents and Multi-channels with Camera-LiDAR Sensors

Zhen Yan, Huijie Fan, Qiang Wang, Pengrui Huang, Yandong Tang · 2025

LiDAR and camera are two different sensors in autonomous driving system. LiDAR can provide accurate object localization and camera can provide rich semantic information. Fusing LiDAR and camera information is essential for 3D object detection. How-ever, it's a significant challenge that how to combine the LiDAR feature and camera feature efficiently. Existing top-performance$3D$object detectors usually rely on the multi-modal fusion strategy within Bird's Eye View representation space. Typically, these 3D object detectors fuse multi-modal features in Encoder modal and then predict in Decoder modal. However, this design has fundamental shortcoming. Due to the inherent differences in the representation of single modal, any choice of modality for expressing mixed features can lead to severe information loss in the other modality. This ultimately restricts the overall performance of the network. In this paper, the authors introduce a robust 3D detector called MRMCFusion. In this network, two independent feature channels are introduced, each using different representation and being maintained separately. On the other hand, to solve the problem of occlusion in single-frame image, cross-temporal image feature fusion is introduced. Experiments on the large-scale nuScenes dataset show that our proposed method surpasses all prior arts.

Read the paper · More papers on PaperTik