Modal-Adaptive Spatial-Aware-Fusion and Propagation Network for Multimodal Vision Crowd Counting
Kai Liu, Xuetao Zou, Peng Zhu, Jun Sang · IEEE Transactions on Consumer Electronics · 2025
With the rapid development of AI and deep learning technologies, smart consumer electronic devices integrate a variety of sensors, especially thermal imaging or depth imaging technology. Some researchers became keen to explore multimodal vision crowd counting techniques based on RGB-thermal or RGB-depth images. However, existing multimodal counting methods ignore the unique contribution of each modality and the effective propagation of features while achieving multimodal fusion. To address these issues, we propose a Modal-adaptive Spatial-aware-fusion and Propagation Network (MSPNet) for multimodal crowd counting. First, we design a coarse-fine cross-modal dynamic fusion (CCDF) module, which fully captures modality-specific and modality-shared information in different semantic spaces using a spatial activation function-guided fine-grained adaptive weighted fusion mechanism. Then, we adopt adapted different convolutions to channel concatenate the features output by multi-stage CCDF. Finally, considering the visual differences of the same object under different imaging angles, we design a dual-dilation adaptive feed-forward propagation module constructed by the dual dilation transposed attention and the feed-forward neural network improved by the gating mechanism to effectively propagate the enhanced multi-modal fusion features to the counting regression module. Experimental results show that MSPNet outperforms the existing state-of-the-art methods on two multimodal crowd counting datasets ShanghaiTechRGBD and RGBT-CC.