Object Tracking and Re-identification in Multi-Camera Environments
Atif Ilyas · HAL (Le Centre pour la Communication Scientifique Directe) · 2011
The video surveillance domain shows very strong growth in recent years. But the proliferation of cameras in public or private spaces makes it extremely difficult for human operators to analyze the data produced by these systems. Many techniques for automatic analysis of the video have been proposed by researchers, and begin to be commercially available. But most of these systems consider the cameras independently one of each other. The objective of this thesis is to address the wide area surveillance, covered by multiple non-overlapping field of view cameras. One of the problems we are interested in is the objects re-identification: when an object appears in the field of a camera, the system should decide whether this object has already been observed and monitored by one camera system or it is a new object. We want to perform this task without any a priori knowledge of the cameras position relative to each other. In the literature, many algorithms exist for moving objects tracking in a video. These algorithms are sufficient to detect object trajectories and to verify that objects have a coherent motion. But these algorithms are not sufficiently robust to object occlusions, intersections, merges and splits. This drawback of current algorithms is problematic, since they form the building blocks of a multi-camera environment. Therefore, the first part of this thesis is to improve the segmentation and object tracking algorithms. At first, we propose an improvement to the foreground/background segmentation algorithms based on codebook. We also propose an evaluation methodology to objectively compare segmentation techniques, based on the analysis of the precision and recall of algorithms. Based on a test set derived from public databases, we show the good behavior of our modified algorithm. A second contribution of this thesis concerns the development of a robust and compact descriptor for moving object tracking in videos. We propose a simple 1-D appearance model, called the Vertical Feature (VF), independent of the view angle and of the apparent size of objects. This descriptor provides a good compromise between very compact color models, that lose all the spatial information of tracked object’s color, and traditional appearance models, too expensive for deformable objects. We associate a motion model of tracked objects and our descriptor, and show the superiority of a combined model approach on traditional tracking approaches, based on the mean shift or on Kalman filter. A descriptor is associated with each object tracked by a camera. Multi-camera tracking, we presents a variability of these descriptors, due to changes in lighting conditions, and also due to the technical characteristics of the cameras, which can differ from one model to the other. We are therefore interested in the problem of the cameras color calibration in order to make similar the descriptors of a same object, seen by different cameras in the system. Existing approaches estimate the Brightness Transfer Functions (BTF) by measuring the response of each camera using known objects. We compare methods based on the Mean BTF (MBTF) and on Cumulative BTF (CBTF) of their color histograms, and show the weaknesses of these approaches when some colors are not enough represented in the objects used for calibration. We propose an alternative (MCBTF) algorithm and we show its superiority over existing methods. Finally, systematic experiments are conducted on the objects re-identification problem in a multi-camera environment, which allows validating all of our proposed algorithms.