A Lightweight Multimodal Fusion Method for Object Detection Based on Bird’s Eye View

Hongmei Chen, Suhong Chang, Wen Ye, Dongbing Gu, Qiangwei Xu, Miaoxin Ji · IEEE Sensors Journal · 2025

This paper proposes a lightweight multi-modal fusion object detection algorithm based on Bird’s Eye View (BEV) perception, addressing the challenges of high computational complexity and insufficient feature fusion when integrating camera and LiDAR data in autonomous vehicles. To alleviate the computational burden, depthwise separable convolution and large separable kernel attention are introduced, constructing a lightweight feature extraction network that effectively captures features from camera data. To enhance computational efficiency, a parallel-optimized BEV pooling structure is proposed that improves the computation process and memory access patterns. In the feature fusion stage, a novel dual-modality dual-attention feature fusion module is designed that integrates the features from both modalities using parallel channel attention and spatial attention mechanisms to strengthen correlations between multi-modal features. Additionally, a weight generation network is introduced to adaptively assign fusion weights to features of the two modalities. The proposed algorithm has an lightweight structure and achieves cross-modal feature alignment in the BEV space.The experimental results of the public nuScenes dataset show that the proposed algorithm achieves an mAP of 0.682 and an NDS of 0.710, while maintaining detection accuracy similar to the baseline model. Furthermore, the algorithm reduces the computational load from 253.2 G MACs to 202.3 G MACs, approximately a 20% decrease, and improves the inference speed from 8.4 FPS to 10.0 FPS.Furthermore, our algorithm has also been experimentally validated on the Waymo dataset, demonstrating performance comparable to that of the BEVFusion method.These results demonstrate significant improvements in both computational efficiency and deployment-friendliness while preserving detection performance.

Read the paper · More papers on PaperTik