Counting Pedestrians in Crowds Using Viewpoint Invariant Training
Dehui Kong, Douglas Andrew Gray, H. Tao · 2005
This paper describes a learning-based method for counting people in crowds from a single camera. Our method takes into account feature normalization to deal with perspective projection and different camera orientation. Thus, our system is trained to be viewpoint invariant and can be deployed with minimal setup for a new site. This is achieved by applying background subtraction and edge detection to each frame and extracting edge orientation and blob size histograms as features. A homography is computed between the ground plane and the image plane coordinates for the region of interest (ROI). A density map that measures the relative size of individuals and a global scale measuring camera orientation are also estimated and used for feature normalization. The relationship between the feature histograms and the number of pedestrians in the crowds is learned from labeled training data. The two training methods used in the current system are linear fitting and neural networks. Experimental results from different sites with different camera orientation demonstrate the performance and the potential of our method. 1