DBF-Mamba: A Mamba-based dual-branch model for multi-modality fusion

Yuxi Luo, Hongyuan Yue, Zhen Xu, Xin Peng, Ziyang Xiao, Luming Li, Hua Wang, Zhiping Wu · IET conference proceedings. · 2025

Multi-modal fusion aims to integrate information from heterogeneous modalities to achieve more accurate and robust scene perception. In recent years, advances in deep neural networks have significantly improved the performance of multi-modal perception in fields such as autonomous driving and medical diagnostics. However, existing methods exhibit limited capacity to model global dependencies and perform efficient multi-modal integration in perception scenarios. Traditional convolutional neural networks (CNNs) are constrained by local inductive bias, enabling them to be less effective at modelling global dependencies. Transformer-based architectures suffer from computational complexity due to the self-attention mechanism, which limits their efficiency and scalability in 3D perception tasks. To address this issue, we propose a Mamba-based multi- modal fusion framework, which enables efficient and scalable multi-modal feature integration. Firstly, a dual-branch encoder is introduced to extract modality-specific features from images and point clouds. Then, a Mamba fusion module is proposed to combine complementary information from different modalities. Experimental results demonstrate that our algorithm outperforms M-CONet by 5.5% IoU and 5.7% mIoU. Compared to existing approaches, our method reduces memory usage during training by about 41.3% and overall inference time by about 34.5%, which validates its robustness and accuracy in complex environments.

Read the paper · More papers on PaperTik