A Multi-Stage Cross-Modal Fusion Framework for Visible–Infrared Object Detection
Anfu Zhu, Yinbing Chen, Heng Guo, Zhizeng Zhang, Yaning Yang, Qinghua Jiang, Yueyong Li, Yi Yang · Applied Sciences · 2026
In visible–infrared object detection under complex environments, cross-modal fusion often suffers from spatial misalignment, semantic inconsistency, and unstable feature responses under varying illumination conditions. To address these issues, this paper proposes a dual-branch visible–infrared object detection framework based on YOLOv11. A staged refine–interact–modulate cross-modal fusion structure (RIFN) is designed to progressively stabilize the cross-modal fusion process. Specifically, the Cross-Modal Spatial–Channel Calibration (CSK) module first suppresses cross-modal discrepancies, the bidirectional Cross-Modal Synergy (CSM) module subsequently performs complementary interaction modeling, and the Complementary Modulation (CM) module further recalibrates fused representations to improve fusion stability under complex environments. Extensive experiments conducted on the public LLVIP and M3FD datasets demonstrate the effectiveness and generalization capability of the proposed framework. Compared with DEYOLO, the proposed method improves mAP@50 by 2.4% on the M3FD dataset while reducing GFLOPs by 2.4, demonstrating superior computational efficiency. In addition, compared with the Transformer-based GM-DETR, the proposed framework improves mAP@50 by 2.1% on the LLVIP dataset and increases inference speed by 210 FPS, achieving higher detection accuracy together with better inference efficiency. Experimental results demonstrate that the proposed staged cross-modal fusion strategy can effectively enhance the robustness of visible–infrared object detection under complex illumination conditions and background interference.