A Multi-Scale Spatial-Temporal Attention Model for Person Re-Identification in Videos
Wei Zhang, Xuanyu He, Xiaodong Yu, Weizhi Lu, Zheng-Jun Zha, Qi Tian · IEEE Transactions on Image Processing · 2019
In this paper, we propose a novel deep neural network based attention model to learn the representative local regions from a video sequence for person re-identification. Specifically, we propose a multi-scale spatial-temporal attention (MSTA) model to measure the regions of each frame in different scales from the perspective of whole video sequence. Compared to traditional temporal attention models, MSTA focuses on exploiting the importance of local regions of each frame to the whole video representation in both spatial and temporal domains. A new training strategy is designed for the proposed model by incorporating the image-to-image mode with the videoto- video mode. Extensive experiments on benchmark datasets demonstrate the superiority of the proposed model over state-ofthe- art methods.