P2SNet: Can an Image Match a Video for Person Re-Identification in an End-to-End Way?

Guangcong Wang, Jianhuang Lai, Xiaohua Xie · IEEE Transactions on Circuits and Systems for Video Technology · 2017

We address a new person re-identification problem, in which our goal is to directly match image-based scenarios with video-based ones. This differs significantly from the conventional person re-identification problem, which aims to match two image-based scenarios (and it is assumed that the available video frames have been manually selected to form the image-based scenarios). To solve this more challenging and realistic problem without the implicit assumption of manual selection, we propose an end-to-end matching framework called a point-to-set network (P2SNet), which consists of: 1) a k-nearest neighbor triplet module, which functions as a “denoiser” by letting the network sequentially focus on the available frames, while ignoring the other useless frames in a video and 2) a novel deep neural network that uses videos and images as input to jointly learn the feature representations and a point-to-set distance metric in a unified way. Our P2SNet is evaluated on three new image-to-video person re-identification data sets, i-LIDS-VID-P2S, PRID2011-P2S, and MARS-P2S, which are modified from i-LIDS-VID, PRID 2011, and MARS, respectively. The experimental results demonstrate the superior performance of our model over the other state-of-the-art methods.

Read the paper · More papers on PaperTik