MSFNet3D: Monocular 3D Object Detection via Dual-Branch Depth-Consistent Fusion and Semantic-Guided Point Cloud Refinement
Rong Qian Yang, Zhijie You, Renhui Luo · World Electric Vehicle Journal · 2025
The rapid development of autonomous driving has underscored the pivotal role of 3D perception. Monocular 3D object detection, as a cost-effective alternative to expensive lidar systems, is garnering increasing attention. However, existing pseudo-lidar methods encounter challenges such as coarse quality and insufficient semantic information when generating 3D point clouds from monocular images. To address these issues, this paper introduces MSFNet3D, which aims to overcome the quality limitations of pseudo-lidar point cloud. Our contributions are threefold: (1) We introduce a dual-branch network to optimize depth maps and propose a multi-scale channel spatial attention module (MS_CBAM). This module captures multi-scale geometric features through a hierarchical feature pyramid and an adaptive weight allocation mechanism, thereby addressing the scale sensitivity inherent in traditional attention mechanisms. (2) We propose a consistency-weighted fusion strategy that employs local gradient consistency analysis and differentiable weighted optimization to achieve a pixel-level fusion of image and depth features. This approach reduces feature conflicts within the dual-branch network and enhances the model’s robustness in complex scenes. (3) We introduce a semantic-guided pseudo-point cloud enhancement method that leverages an instance segmentation network to extract object-specific semantic regions and generate high-confidence point cloud, consequently improving the accuracy of object detection. Experiments on the KITTI dataset show that the proposed method performs excellently under various detection challenges, achieving an average precision of 18.87% in the 3D detection of car objects, which is a 1.67% improvement over the original model. The method also shows good performance in detecting pedestrians and cyclists. The proposed framework can provide economical and reliable 3D perception for mass-produced electric vehicles.