Text Self-Supervision Enhances Visual-Language Learning for Person Re-Identification

Bin Wang, Huakun Huang, Guowei Liu, Yatie Xiao, Lingjun Zhao, Zhenbang Liu · 2024

Recently, some visual-language learning-based methods have overcome the lack of text descriptions in the person ReID. By introducing large-scale vision-language pre-trained models like CLIP, these methods have shown impressive performance in person ReID tasks. However, those methods rely on fixed parameters of pre-trained CLIP to generate text descriptions for ReID data via visual-language learning. As CLIP is not specifically trained and tailored for ReID, generated text descriptions limited the model’s fine-grained feature extraction capability. Therefore, we aim to generate more discriminative text descriptions by incorporating text self-supervision. To achieve this, we propose using a PCA module combined with text contrastive learning for text self-supervision alongside the original visual-language learning. We conduct experiments on four commonly used ReID datasets. Our method adapts the model to the feature distribution of this task and improves its generalization ability. The proposed TSLIP-ReID outperforms leading state-of-the-art methods.

Read the paper · More papers on PaperTik