Steam: Sparse Transformer and Explicit Attention Module for Multimodal Object Detection

Xiaoxiong Lan, Shenghao Liu, Sude Zhang, Zhiyong Zhang · IEEE Sensors Journal · 2025

Multimodal object detection by combining multi-sensor data through visible-thermal image pairs has shown significant potential in complex environments. However, the information redundancy and modality conflicts between these modalities limit further performance improvements. Existing methods often fuse features through direct input into CNNs or Transformers, but these approaches fail to sufficiently address the challenges posed by the distinct imaging mechanisms of the modalities. To overcome these limitations, we introduce a novel Steam module, which incorporates two key innovations: the Sparse Attention Module (SAM) and the Explicit Attention Module (EAM). The SAM module extracts local information using CNNs and captures global information while eliminating redundancy through an improved sparse Transformer. The EAM module resolves modality conflicts through an explicit attention mechanism that makes sure to activate the target feature. Based on this module, we propose a new two-stream one-stage multimodal object detection framework named SteamDet. Meanwhile, to reduce the amount of computation, we downsample the feature maps before the fusion module and design an Efficient Upsampling Block (EUB) to recover. Extensive experiments demonstrate that our method achieves state-of-the-art performance across three public datasets, improving mAP by 6.1% on the multimodal-oriented DroneVehicle dataset and achieving optimal results across all categories. Furthermore, on the FLIR and LLVIP datasets, mAP increases by 2.3% and 2.9%, respectively, highlighting the strong generalization ability of our approach in addressing redundant information and modality conflicts. Our code will be released at https://github.com/lanxx314/steam.

Read the paper · More papers on PaperTik