Visible-Infrared Person Re-Identification via Mutual Reinforcement of Prompts and Image Encoders

Hongde Zhang, Bingpeng Ma · 2025

Contrastive Language-Image Pre-training (CLIP) has achieved good results in Visible-Infrared Person Re-IDentification (VI-ReID) task. However, CLIP does not focus on person-related information, so prompts generated by original CLIP can not accurately describe identity information of a person. We argue that compared to original CLIP, encoders familiar with person-related information can generate prompts which are more suitable for VI-ReID. Based on such idea, we design a novel network that helps prompts focus on person-related information through alternately optimizing the prompts and image encoders. Specifically, when optimizing prompts, we introduce modality knowledge propagation loss. The loss aligns the predicted class probability of text and image features, so that the knowledge in image encoders is transferred to prompts. When optimizing encoders, we design modality alignment loss. The loss considers text features as a bridge between two modalities, aligning features from two modalities with text features. In this way, modality discrepancies are effectively reduced. Finally, through the mutual reinforcement of two parts, the quality of both prompts and image encoders is improved in a positive feedback manner. Experiments on two widely used datasets show that the proposed network outperforms state-of-the-art methods.

Read the paper · More papers on PaperTik