Mamba-driven cross-modal state-space fusion: a dynamic framework for robust pedestrian detection
Jiawei Fan, Dan Wei, Xiaolan Wang · Journal of Electronic Imaging · 2025
Most of the existing modal fusion methods for multimodal pedestrian detection rely on the traditional neural network architecture, which is limited by the inherent local reduction bias or the quadratic computational complexity, and it cannot fully capture the interaction between the modes. However, the latest research showed that the Mamba architecture based on the state-space model achieves efficient feature extraction with linear computational complexity in long-sequence modeling through a selective scanning mechanism and hardware-aware optimization. Mamba shows higher parameter utilization and linear memory growth characteristics than Transformer in the visual-verbal-temporal cross-modal task. Therefore, we propose a multimodal fusion pedestrian detection scheme based on Mamba. Cross-modal fusion is investigated by associating cross-modal features in a modified Mamba-based hidden state space with a gating mechanism. We design a cross-modality fusion Mamba block to map the modal features of infrared images with those of visible images to a hidden state space for interaction, thus reducing the differences among different modal features. A dual-state-space channel exchange module is designed to promote the fusion of shallow features. A dynamic two-state-spatial fusion module is designed to realize the dynamic deep fusion in the hidden state space. The contribution of different modal features is adaptively adjusted by the gating mechanism to further reduce the modal difference, improving single-mode detection performance. The effectiveness of our model is verified by extensive experiments on four widely used visible-infrared benchmark datasets. The results show that M-fusion has excellent performance in multimodal pedestrian detection.