OIGD-MDFusion: A Multi-Modal 3D Object Detection Framework Combining Internal Geometry Distillation and Decoupled Attention
Zhipeng Shi, Yuanyao Lu, Haowei Yang · 2025
With the rapid advancement of autonomous driving technology, the demand for efficient and accurate 3D object detection systems has significantly increased. To achieve effective cross-modal 3D object detection, existing methods often enhance LiDAR-based detection performance by leveraging the rich semantic information provided by the camera modality, such as improving depth perception and semantic understanding through image branches. However, these approaches typically rely on simplistic feature fusion strategies to align camera and LiDAR features, failing to fully account for the geometric and semantic differences between the two modalities. This limitation adversely impacts detection accuracy and system robustness. To address this issue, this paper proposes an innovative multimodal detection framework, OIGD-MDFusion. The framework incorporates an Object-Intrinsic Geometry Distillation (OIGD) module to improve the alignment accuracy between camera and LiDAR features and employs a Multi-Decoupled Attention Module to enhance the expressive capabilities of both point cloud and image modalities. Experimental results on the nuScenes benchmark demonstrate that OIGD-MDFusion achieves significant improvements in detection accuracy while maintaining outstanding computational efficiency. Notably, compared to existing methods, our model exhibits strong competitiveness on the nuScenes test benchmark, achieving a 72.9% mAP and 75.6% NDS.