Multi-Modal BEV Enhancement Fusion for 3D Object Detection in Autonomous Driving

Chen Mu, Mingchuan Yang, Yuan Zhang, Tao Han, Xinchi Li, Huaici Zhao, Pengfei Liu · IEEE Transactions on Intelligent Transportation Systems · 2025

Recent success in 3D object detection have underscored its importance in autonomous driving, particularly through the integration of diverse sensor modalities like RGB images and LiDAR point clouds. With the benefits of Lift-Splat Shift (LSS) paradigm, different data modalities can be effectively fused in Bird’s-Eye-View (BEV), significantly improving the detection performance. Although BEV-based fusion has significantly advanced 3D detection technology, the limited enhancement of image features and the inconsistency between different modalities still hinder the overall performance of the detector. In this work, we focus on effective enhancement strategies and design 3D object detection pipeline named ECL3D to further push the detection performance boundary. The first strategy, Depth-Semantic Feature Enhancement (DSE), aims to improve input features for the view transformer during training without adding computational burden during inference. This approach leverages low-resolution depth distribution supervision to maintain the accuracy of frustum generation and high-resolution depth supervision to provide richer clues for front-end features. The second strategy, Instance BEV Feature Enhancement (IBFE), introduces a mechanism to enhance instance-relevant features for multi-modal fusion. This strategy effectively suppresses background noise and enhances object-region features. Comprehensive experiments on nuScenes dataset demonstrate the effectiveness of our approach. Without any test-time-augmentation strategy, our detector achieves state-of-the-art performance in camera-LiDAR fusion 3D object detection task with mAP and NDS of 72.8% and 75.2%, respectively. The code is coming soon at: https://github.com/muchen2019/ECL3D

Read the paper · More papers on PaperTik