CMIF-YOLOv5: An Object Detection Algorithm Based on Cross-Modal Interaction Fusion

Guodong Zhang, Hongran Zhao · 2025

In adverse environments such as night, fog, and glare, mainstream unimodal single-modal object detection models exhibit significant performance degradation. Therefore, scholars have explored the fusion of visible and infrared images to improve object detection performance. However, adverse environments can interfere with sensors, introducing noise and irrelevant information into the images. Existing modality fusion methods often overlook these noise factors, and the large discrepancies between different modality features can adversely affect detection performance when fused directly. To address this, we propose an object detection algorithm, CMIF-YOLOv5, based on modality interaction attention module. Using YOLOv5 as the baseline model, a dual-stream feature extraction backbone network is employed to replace the original backbone network. This can overcome the limitations of single-modal features. The modal interaction attention module was designed. Mitigating differences between different modal features in channel dimension and spatial dimension. The SIoU loss function is used in the detection head to optimize the IoU between the predicted box and the ground truth box.

Read the paper · More papers on PaperTik