Learning Modality Advantage Hierarchically for RGB‐Infrared Object Detection
Zelin Xiang, X.C. Li · IET Image Processing · 2026
ABSTRACT The core of multi‐modal object detection lies in effectively utilizing the complementary strengths of different modalities to enhance performance. However, existing methods often focus on fusion strategies while neglecting the fact that different targets exhibit distinct advantages in specific modalities. For example, pedestrians are more distinguishable in infrared images, while vehicles show clearer textures in RGB images. Without explicitly modeling these modality‐specific advantages, complementary information remains underutilized, limiting detection efficacy. This paper introduces a hierarchical modality advantage learning (HMAL) framework to address this issue by transforming modality advantages into learning guidance signals and enhancing feature fusion across multiple levels. At the low level, a triple‐stream collaborative detection encoder is developed with a dual‐stream structure to preserve RGB and infrared‐specific features. An auxiliary modality‐aware fusion unit is also introduced to align low‐level details. At the high level, a modality advantage guided learning module quantifies each target's modality advantage and injects this information into semantic features through learnable embeddings. By combining advantage‐guided semantics with low‐level details using cross‐attention, the framework achieves comprehensive enhancement. This hierarchical approach leverages the unique strengths of each modality while improving cross‐modal and cross‐level information integration, significantly enhancing detection performance. The source code and are available at echo9958/HMAL.