Motion pattern learning and applications to tracking and detection
Xuemei Zhao · University of Southern California Digital Library · 2013
With the decreasing cost of collecting data, the deluge of surveillance videos makes it necessary to carry out automatic intelligent processing to understand scenes and analyze activities. There are two classes of methods to video analysis, one is based on analyzing the trajectories of objects of interest, and the other is examining motion features vectors directly. Our work is a combination of the two. ? Learning motion patterns of moving objects, a key problem in the second category, is an important way for scene understanding, since motion patterns convey rich information of the scene structures. In this dissertation, we first develop an unsupervised learning framework to infer motion patterns in videos, and in turn use them to assist tracking and detection of objects, especially in crowded scenes. ? Based on motion features such as optical flow or tracklets input, we embed feature points into motion feature space, and use a manifold learning method Tensor Voting to infer the local geometric structures. In this space, points exhibit intrinsic manifold structures, each of which corresponds to a motion pattern. To define each group, a novel robust manifold grouping algorithm is proposed. Tensor Voting is performed to provide multiple geometric cues, such as local tangent/normal space and dimensionalities of the geometric structures that a point belongs to. Multiple similarity kernels based on the geometric cues are formulated between any pair of points, and a spectral clustering technique is used in this multiple kernel setting. A generalization of the grouping algorithm to N-D space is proposed and tested on various dataset, including synthetic data, USPS digits, CMU Motion Capture data (MoCap) and real-world video sequences. It achieves better performance than state-of-the-art methods in these applications. Furthermore, an online distributed framework of motion pattern learning is proposed to deal with big data, such as wide area aerial surveillance imagery. ? To understand scenes, tracking and detection of general objects are necessary and challenging tasks. In our work, we focus on solving the problem in structured scenes, i.e., the scenes with clear motion patterns. The most salient characteristic of structured scenes is, objects do not move randomly, but follow some patterns instead. Extracted motion patterns convey rich information such as how the objects move, and how they interact with each other. Using them as a prior, we significantly improve the performance of tracking and detection. ? In video analysis, a large group of videos which are both interesting and challenging is high density crowded video scenes containing hundreds of similar objects, such as heavy traffic roads, busy shopping malls, political gatherings, sports events, etc. These scenes need special attention for the purpose of public safety and accident prevention. However, tracking and detection in crowded scenes are very challenging due to the large amount of similar objects, cluttered background, small target size, and occlusions caused by the constant interactions between objects. In our work, we investigate the single and multiple target tracking and detection problems in structured crowded scenes, and propose Motion Structure Tracker (MST) based on motion pattern information. It is a combination of motion pattern learning, visual tracking and multi-target tracking. In MST, tracking and detection are performed jointly, and motion pattern information is integrated in both steps to enforce scene structure constraint. Experiments are performed on real-world challenging sequences, and MST gives promising results. The comparison to several state-of-the-art methods proves its effectiveness.