Exploitation of wide area motion imagery
Jan Prokaj · University of Southern California Digital Library · 2013
Current digital photography solutions now routinely allow the capture of tens of megapixels of data at 2 frames per second. At these resolutions, a geographic area covering a whole city can be captured at once from an unmanned aerial vehicle (UAV), while still allowing the recognition of vehicles and people (for sensors under development). This fact, in tandem with the availability of increased computational power, has led to the growth of wide area motion imagery (WAMI). ? This imagery opens doors to a lot of applications, for example, in urban planning, security, and geospatial digital libraries. However, we have found out that this imagery cannot be easily processed by human operators, just because of its sheer amount. Therefore, we propose that a paradigm shift is required to enable working with UAV video data. The task of visualization cannot be simply defined as looking at the imagery. In reality, the pixel stream is a signal rich with information that must be extracted into a form that is easily processed by a person. The development of this extraction process falls into the realm of computer vision and is the heart of this dissertation. ? WAMI data is often captured by an array of cameras. Therefore, at the lowest level, we need an algorithm that takes an array of individual camera images and estimates a high quality mosaic. We propose a piecewise affine model to handle all image deformations that deviate from the standard pinhole camera model. The results show our model produces better mosaics than a standard lens distortion model. ? Detection of moving objects is not possible with a moving camera. Video stabilization transforms the mosaicked video stream into one without any camera motion. We propose a stabilization technique which minimizes the amount of drift and jitter in the stabilized imagery. The results show our technique has better accuracy compared to widely used techniques. ? The next level of processing involves estimating the trajectories of all moving objects, or "tracking." We propose a tracking algorithm that optimally infers short tracks using Bayesian networks. These tracklets are then integrated into a multi-object tracking algorithm that achieves good performance on aerial surveillance video. When coupled with a regression-based tracker, stopping targets can be handled. ? WAMI is often collected over urban areas, where there are tall buildings, and other structures, which cause severe occlusion that in turn causes significant track fragmentation. To solve this problem, we propose a method which links fragmented tracks using known 3D scene structure. Our method outperforms the classic Hungarian algorithm. ? In order to enable large scale semantic analysis of WAMI data, higher level algorithms that determine at least some of the semantics are necessary. Expecting full semantic description of the scene is unrealistic, but the task can be made easier for a human operator by automatically determining the most common, or primitive, events or activities. We propose a framework based on the Entity Relationship Model that is able to recognize a large variety of activities on real data as well as GPS tracks.