DIRL: Learning Discriminative ID-Related Representations for Video Visible-Infrared Person ReID

Jiahe Wang, Xizhan Gao, Sijie Niu, Hui Ying Zhao, Guang Feng · ACM Transactions on Multimedia Computing Communications and Applications · 2025

The core of Video-Based Visible-Infrared Person Re-Identification (VVI-ReID) lies in learning modal-sharing features, namely ID-related feature representations, that are often mixed within modal-invariant features. Existing methods use weight sharing networks to learn frame-level modal-invariant features, but fail to consider modality invariance when computing sequence-level features, resulting in a significant gap between modalities. Moreover, these methods do not explicitly separate the ID-related features and modal-related features, which reduces the model’s discriminative ability and leads to interference in VVI-ReID. In this article, we propose a Discriminative ID-Related Representation Learning (DIRL) network for VVI-ReID. DIRL network consists of three key components, that is Two-Stream Backbone Module (TBM), Cross-Modality Interaction Module (CIM), and Feature Decoupling Module (FDM). More specifically, the TBM is first constructed to preliminarily capture frame-level modal-invariant features. Then, the CIM is designed to interact information between modals and aggregate temporal features simultaneously, thereby obtaining sequence-level modal-invariant features. Finally, the FDM is designed to explicitly separate modal-related features from ID-related ones within the modal-invariant features, thereby leaving only discriminative ID-related representations. Through extensive benchmark experiments, our method demonstrates superior performance over state-of-the-art approaches by significant margins. Our code will be available at https://github.com/JhSearch/DIRL .

Read the paper · More papers on PaperTik