3D Object Detection Method With Cross-Modal Attention Mechanism for LiDAR-Camera Feature Fusion
Quanbo Yuan, Xu Liu, Xinyue Chen, Huijuan Wang, Peng Liu, Guocai Yin, Shilong Jin, Jianhua Wang · International Journal of Image and Graphics · 2026
With the advancement of computer vision technology, 3D object detection has emerged as a pivotal technique in autonomous driving, robotics, and virtual reality. Multimodal fusion approaches often outperform single-modal methods in 3D object detection, yet they face challenges such as feature redundancy, insufficient fusion, and loss of spatial resolution. To address these issues, we propose HAM-Net, a hybrid attention-based multimodal 3D object detection method. HAM-Net employs an SEFusion module to adaptively weight and fuse features from LiDAR and camera data, enhancing the model’s ability to focus on critical features. A hybrid attention module, integrating channel, spatial, and BAM attentions, further optimizes feature selection. Additionally, we improve the FPN module to restore feature map resolution, mitigating information loss. Experiments on the KITTI dataset demonstrate that HAM-Net significantly outperforms baseline models, achieving 87.84% and 78.48% mAP in aerial view and 3D detection, respectively. Our contributions lie in the novel SEFusion strategy, hybrid attention mechanism, and optimized FPN, collectively enhancing the performance, generalization, and adaptability of multimodal 3D object detection.