Human detection in images and videos

Kwang-Ting Tim Cheng, Qiang Zhu · 2007

This dissertation addresses the problems of detecting objects in images and videos with a particular emphasis on human detection. Object detection and recognition is a critical task for many real-world applications. Recent studies have resulted in several general approaches that combine salient local image features and effective learning techniques. We developed near real-time systems for automatically finding people under two practical scenarios: (I) Under the assumption that humans can be found in isolation (e.g., in personal photos), our solution integrates a cascade-of-rejector approach with the Histograms of Oriented Gradients (HoG) features to achieve a fast and accurate human detection system. The features used in our system are HoGs of variable-size blocks that capture salient features of humans automatically. (II) For the situations where humans move in crowds (e.g., the exit of a railway station), we present a system which first detects interest points in each frame, tracks them over time and clusters them. The key challenge is to cluster the feature points that are in a high dimensional, non-Euclidean space and are of variable length, depending on the number of frames they appeared in so far. In both approaches, we explore the effectiveness of using local image features for representing an object, and then employing machine learning techniques for detecting the target object in images and videos. Extensive experiments on real data demonstrate that the proposed approaches achieve high accuracy in practice.

Read the paper · More papers on PaperTik