GLFNet: An RGB-T Crowd Counting Network Based on Global–Local Multimodal Feature Fusion

Yingxiang Hu, Yanbo Liu, Guo Cao, Jin Wang · IEEE Transactions on Instrumentation and Measurement · 2025

RGB-T crowd counting methods aim to enhance the counting accuracy of network models under conditions of uneven lighting and low visibility by fusing features from the RGB and thermal modalities. Previous approaches primarily utilized attention mechanisms to extract and fuse complementary RGB and thermal features. However, these methods lack guidance and constraints during the extraction and fusion of multi-modal features and do not fully leverage the complementary advantages between global and local features, leading to suboptimal performance. This paper argues that, by transitioning from global attention to local attention, extracting and fusing the complementary information between global and local multi-modal features can significantly improve the model’s counting performance. To achieve this, we propose an RGB-T crowd counting network based on global-local multimodal feature fusion (GLFNet). Specifically, we first use a multi-head attention mechanism to fuse global multi-modal features and guide the global multi-modal fusion using learnable block-counting guided tokens (BCT). Next, we employ composite spatial attention mechanisms (CSAM) to focus on the local detail information of multi-modal crowd features and facilitate the fusion of local multimodal features. Finally, we utilize a detail contrast loss function (Ld) to capture the complementary advantages between global and local multi-modal features and to guide and constrain the fusion process of multi-modal features. Experimental results on the RGBT-CC and DroneRGBT datasets demonstrate the superior performance of our method.

Read the paper · More papers on PaperTik