A Comprehensive Deep Learning Model for Improved Person Re-identification Using Multi-Camera Streaming Pipeline
Patel Devarshi, Chauhan Yash, Meghna B. Patel, Dhaval K. Raval · Procedia Computer Science · 2025
The growing importance of multi-camera person Re-ID reflects its significance in computer vision. While single-camera tracking has advanced, maintaining identification across multiple cameras remains challenging due to factors like occlusion, appearance changes, camera motion, and lighting conditions. In this proposed research work, present a comprehensive multi-camera person Re-ID process designed to address these challenges by leveraging advanced detection and tracking methodologies. The process begins with the utilization of the state-of-the-art YOLOv8n and lightweight object detection model, specifically for person detection. This model efficiently identifies and localizes individuals within the video frames captured by multiple cameras, each covering a unique perspective within the monitored environment. The system integrates DeepSORT trained on the MARS dataset for multiple object tracking (MOT) to maintain consistent identification across video frames. Following person detection and tracking, a custom transformer model built on the TorchReID framework is employed to analyze and compare bounding boxes of detected persons, ensuring that the identified individuals are accurately associated with existing IDs. The Re-ID process is further enhanced by custom trained OSNet_x1_0 and ResNet50 models. These models, trained on a combination of the Market1501 dataset and the UNI09 custom dataset, extract rich and discriminative feature embeddings. The integration of these models has achieved high accuracy in identifying individuals across multiple camera views. Notably, OSNet_x1_0 achieves 98.4% mAP and ResNet50 reaches 96.1% mAP on Market1501. The integration of these models into the Re-ID pipeline has achieved high accuracy in identifying individuals across multiple camera views. Experimental results demonstrate the system’s ability to maintain accurate person identification across cameras, achieving high consistency. This outcome highlights the effectiveness of proposed approach.