SRCFNET: A Query-Based Fully Sparse Network for 4D Radar-Camera 3D Object Detection
Yuran Li, Qiang Ling · 2025
Accurate 3D object detection is critical for autonomous driving systems. Radar-camera fusion has been extensively validated as an effective approach to enhancing detection accuracy, by capitalizing on complementary sensor capabilities. Specifically, the emerging 4D millimeter-wave (mmWave) radar provides precise geometric, spatial, and velocity information with robust performance in adverse weather conditions, while cameras offer rich semantic and texture details. In this paper, we propose the Sparse Radar-Camera Fusion Network (SRCFNet), a lightweight and fully sparse query-based framework designed for 4D radar-camera fusion in 3D object detection. By leveraging predefined queries in the Bird's-Eye View (BEV) space, SRCFNet integrates radar point clouds and image features through our proposed Channel-Adaptive Feature Fusion (CAFF) module, followed by iterative refinement via a transformer decoder. Experiments on the TJ4DRadSet dataset demonstrate significant performance improvements, achieving 3.08% and 2.72% increase in 3 D mAP and BEV mAP, respectively, compared to baseline methods. This lightweight approach provides an effective and scalable solution for autonomous driving perception.