Cross-modal person re-identification based on two-stage feature aggregation and multi-scale refinement enhancement network

Zoufei Zhao, Lihong Li, Qingqing Liu, Ziwei Zeng · Engineering Research Express · 2025

Abstract Cross-modal person re-identification involves matching pedestrian images across visible and infrared modalities. The task is crucial for applications such as criminal investigation. This paper introduces Two-stage Feature Aggregation and Multi-scale Refinement Enhancement Network, a novel network designed to address limitations in fine-grained feature extraction. The network integrates feature information from different stages through a Two-stage Feature Aggregation Module, which combines with a Hybrid Attention Block to effectively capture key features across both the channel and spatial dimensions. Additionally, the Multi-scale Refinement Expansion Block processes multi-scale features utilizing convolutional layers with varying dilation rates, further enhancing the alignment of cross-modal features. To optimize model performance, the paper integrates multiple loss functions to train the network. Comprehensive experiments on the SYSU-MM01 and RegDB datasets show that the proposed approach surpasses the most advanced approaches. Moreover, ablation studies and comparative analyses further validate the superiority and effectiveness of the model. The code for our approach can be accessed at: https://github.com/FF123-hue/TAFN-NeT/tree/master.

Read the paper · More papers on PaperTik