DPSNET: A Dual-Path Lightweight Network With Semantic-Guided Cross-Modal Feature Fusion for Multimodal Object Detection

Yu Jie Huang, Haiyan Li, Guanbo Wang, Dan Xu, Jundong Yang, Geng Yue, Xiang Liu · IEEE Transactions on Instrumentation and Measurement · 2025

Multimodal object detection is crucial for autonomous driving and intelligent surveillance. However, existing methods face critical challenges, such as cross-modal semantic misalignment from heterogeneous spatial structures, degraded infrared representation under low contrast, and inadequate adaptation of visible feature extraction. To fix the gap, we propose a lightweight dual-path network with semantic-guided cross-modal feature fusion for multi-modal object detection. Firstly, to mitigate cross-modal feature-level semantic misalignment, the Semantic-Guided Cross-Modal Fusion (SCMF) module based on a semantic-guided attention is designed, which constructs channel-specific spatial importance maps and enables fine-grained and semantically consistent fusion of heterogeneous features via structured attention and adaptive residual modulation. Secondly, the Adaptive Feature Enhancement Aggregation (AFEA) module is proposed to jointly capture long-range semantic context and fine-grained local structure features by combining global attention inference, local feature estimation, and spatially aware receptive field enhancement. Lastly, the Dynamic Receptive Field Attention (DRAF) module is put forward, which dynamically adjusts feature interactions based on adaptive receptive field modulation to further refine feature-specific fusion, improving robustness under diverse conditions. Experimental results on LLVIP, FLIR and KAIST datasets demonstrate that DPSNET maintains a good balance between detection accuracy and model lightweight, making it a viable solution for multimodal scenarios.

Read the paper · More papers on PaperTik