TFE-Fusion: Tri-Modality Feature Enhanced Fusion for Robust Object Detection

Muhammad Usama, Imad Ali Shah, Roshan George, Enda Ward, Fiachra Collins, William O'Grady, Martin Glavin, Brian Michael Deegan, Edward Jones · IEEE Open Journal of Vehicular Technology · 2026

Robust and reliable object detection under adverse conditions remains a critical challenge for automated driving systems (ADS). The performance of RGB-based (visible-spectrum) cameras degrades in poor lighting conditions, whereas the performance of RGB-Long-Wave Infrared (LWIR) fusion architectures remains limited in adverse weather and low thermal contrast scenarios due to partial spectral coverage. At the same time, Short-Wave Infrared (SWIR) has the potential to address some of these gaps but remains under-utilised in ADS. To address these limitations, we propose a novel tri-modality fusion architecture, TFE-Fusion, that simultaneously leverages information from RGB, SWIR, and LWIR modalities. Our architecture employs a feature-level fusion strategy that incorporates pixel-level weighting and a spectral attention mechanism, enabling dynamic fusion of complementary features from all three modalities. A YOLOv8-based detection head, modified for multi-modal input streams, is used for efficient and robust inference. Extensive experiments on the Multispectral Object Detection (MOD) dataset demonstrate that the proposed method outperforms the RGB-LWIR baseline by 2.48 percentage points in [email protected]:0.95. The most significant performance gains are observed for the vehicle class, where SWIR features compensate for low thermal contrast in LWIR imagery by capturing distinct material reflectance properties. These results validate the effectiveness of the proposed TFE-Fusion architecture for enhancing detection performance, particularly under low-light and low-thermal-contrast conditions.

Read the paper · More papers on PaperTik