Cross-Modal Object Detection Based on Content-Guided Feature Fusion and Self-Calibration

Liyang Ning, Xuxun Liu, Luoyu Zhou, Xueyu Zou · Sensors · 2025

Traditional transformers suffer from limitations in local attention, which can result in inadequate feature representation and reduced detection accuracy in cross-modal object detection tasks. Additionally, deep features are prone to degradation through multiple convolutional layers, leading to the loss of detailed information. To address these issues, we propose a dual-backbone cross-modal object detection model based on YOLOv8n. First, we introduce a parallel network in the backbone to enable the model to process information from different modalities simultaneously. Second, we design a content-guided fusion module (CGF) in the feature extraction network, leveraging both transformer and convolution operations to capture global and local information, thereby enhancing the model's ability to focus on detailed object features. Finally, we propose an adaptive calibration fusion module (ACF) in the neck to fuse shallow and deep features, supplementing fine-grained details and improving the model's detection capability in complex environments. Experimental results show that on the LLVIP dataset, our model achieves mAP50 of 96.4 and mAP95 of 63.8; on the M3FD dataset, it achieves mAP50 of 83.7 and mAP95 of 56.6. Our model outperforms baseline models and other state-of-the-art methods in detection accuracy, demonstrating robust performance for cross-modal object detection tasks across various environments.

Read the paper · More papers on PaperTik