Translation Templates for Object Matching Across Predictable Pose Variation
Chris Stauffer, Matthew Antone · 2006
Computer vision is the most promising technology for automated, passive tracking of multiple objects over large areas. Effective inter-camera and intra-camera visual tracking can enable information about a vehicle or a pedestrian to be integrated from various sources. Unfortunately, tracking objects across multiple non-overlapping cameras requires reliable comparison of the objects ’ appearance under widely-varying view angles and resolutions. Fortunately, in most cases, an object of a particular type entering a scene at a particular position and direction will tend to be in a very similar pose. This paper introduces Translation Templates (TTs). TTs exploit this regularity to learn a color-based matching metric for images from a pair of tracking source and sink points, without prior knowledge of object type or object pose. This model benefits from histogram-based aggregation while still preserving spatial relationships between the two images. The model can be learned directly from data and used to compare arbitrary types of objects observed from extremely different viewpoints, as long as the relationship between the viewpoints is preserved. This paper describes TTs, describes a method for efficient computation and for visualization of TTs, and presents experimental results from both an indoor pedestrian data set and an outdoor vehicle data set. 1