Interlaced Perception for Person Re-identification Based on Swin Transformer
Baofeng Zhang, Yaling Liang, Minghui Du · 2022 7th International Conference on Image, Vision and Computing (ICIVC) · 2022
Extracting discriminative feature presentations is crucial for Person Re-identification (Person Re-ID) tasks. Although convolutional neural networks have powerful feature extraction capabilities, modeling non-local relationships is still challenging for them. A promising strategy is to adopt visual transformer, which excels at modeling long dependencies based on the self-attention mechanism. However, classic visual transformer has shortcomings in focusing on local and fine-grained features. In this paper, we construct a novel and strong Person Re-ID model based on Swin Transformer, which integrates the advantages of both CNN and Transformer. To further enhance the ability to capture long-range dependencies, we introduce Adaptive Split Self-Attention (ASA) module. ASA module improves globally modeling over relations between regions by performing self-attention on the adaptive partitions of the input. Extensive experimental results on several datasets including Market-1501, DukeMTMC, and MSMT17 demonstrate that our method achieves state-of-the-art performance.