DF-Fusion: a multimodal 3D object detection framework based on frequency domain decoupling and dynamic submanifold fusion

Hao Wu, Xiaoqian Shi, Jiawei Li, Qidong Chen · Journal of King Saud University - Computer and Information Sciences · 2026

Multimodal fusion has become the mainstream paradigm for 3D object detection in autonomous driving. However, existing methods still present two challenges in complex outdoor scenarios: Fusion strategies with fixed receptive fields struggle to adapt to uneven spatial scale distributions characterized by high density at close range and low density at a distance, and the weak features of small objects are easily suppressed by the background during downsampling. Cross-modal interaction based on cross-attention tends to aggregate high-frequency background noise under out-of-distribution noise, leading to false positive predictions. To address these issues, this paper proposes DF-Fusion, a multimodal 3D object detection framework driven by frequency domain decoupling and dynamic submanifold fusion. The framework consists of two core modules: depth- and density-aware dynamic submanifold fusion (DSF) convolution, which dynamically adjusts the fusion receptive field of image features on the basis of the voxel depth and local point cloud density to achieve scale-adaptive alignment across all levels of the feature pyramid. The frequency-decoupled latent instance-scene graph (FD-LISG) module uses a discrete wavelet transform to decouple the fused features into low-frequency semantics and high-frequency details and transmit information between instances in the latent graph space using only the low-frequency components while preserving credible high-frequency details via point cloud coordinate masks. Experimental results on the nuScenes dataset demonstrate that among the compared methods, DF-Fusion achieves the highest detection accuracy for seven target classes, including trailer, pedestrian, and bicycle.

Read the paper · More papers on PaperTik