Text-Guided Visual Representation Learning via Cross-Modal Fusion for Person Re-Identification
Ge Cao, Qing Tang, Xuan-Thuy Vo, Adri Priadana, Tien-Dat Tran, Kang-Hyun Jo · IEEE Access · 2025
Person Re-identification (Re-ID) aims at accurately querying pedestrians across multiple non-overlapping cameras system, playing an essential role in computer vision applications. While CNN-based methods leverage feature aggregation and attention mechanisms to achieve competitive performance, their limitations in capturing long-range dependencies have motivated the research of transformer-based architectures, which excel in the evolution of person Re-ID research. Despite leveraging pretrained vision-language models to facilitate visual encoder training, current methods suffer from fragmented optimization procedures, complexity increasing, and inadequate semantic guidance from textual prompts. To address these limitations, we propose an end-to-end method named Text-Guided Fusion Transformer (TGFT). TGFT applies the fixed but semantically enriched texts to guide visual training with a novel Gated Cross Attention (GCA) fusion module. The GCA module adaptively modulates the integration of textual information and visual features, effectively enhancing cross-modal feature alignment and semantic consistency. Without significantly increasing computational complexity or parameter count, TGFT achieves superior performance, demonstrating strong effectiveness and scalability on multiple major Re-ID benchmarks.