Ensemble Learning using Vision Transformer and Convolutional Networks for Person Re-ID

Aryan Kumar Gupta, Neil Gautam, Dinesh Kumar Vishwakarma · 2022 6th International Conference on Computing Methodologies and Communication (ICCMC) · 2022

Person Re-Identification is the process of recognizing a targeted individual across multiple views at different times, in different and challenging real-life diverse settings. It remains a conundrum due to the significant amount of intra-class variation present in same individual caught across different cameras. Most of the existing models require a large amount of data for training, as a result of which they do not generalize well on small datasets and hence decreases the robustness of the identification process. To reduce this variance, this paper introduces an end-to-end triple stream ensemble model making minimal changes in the Vision Transformer, Resnet50 and Densenet121 architectures respectively. Our model performs well on the Market1501 dataset achieving an accuracy of 90.05% and 80.45% on the Duke MTMC ReID dataset.

Read the paper · More papers on PaperTik