Template-based descriptors for object detection and tracking
Stefano Soatta, Taehee Lee · 2012
Detecting and tracking objects is an important problem for several tasks using computer vision. Especially with the help of recently developed and deployed personal mobile devices, such as phones or tablet computers with cameras, it becomes a keystone to performing interactive object based tasks. For the visual search applications, designing a good feature descriptor is important for improving the recognition accuracy under various conditions of query images and target objects. Several nuisances including viewpoint changes, occlusions, and lighting changes make the visual recognition problem more difficult and require the feature descriptors to be invariant or robust to such changes. In addition to recognizing, tracking features and objects are needed for some visual tasks. For example, augmented reality applications require real-time rendering of visual contents aligned to video input. Performing recognition for every frame is not only wasting computation power but also ignoring the temporal consistency of the objects between consecutive frames. By combining tracking and recognition for video, such needed real-time performance can be achieved efficiently and the recognition accuracy can be improved by associating more data from video than from a single image. In this thesis, we first analyze and describe how to learn the best template descriptors from video, and introduce an efficient and robust feature tracking algorithm that can be combined for learning such descriptors. Then we investigate on the case of significant scale changes, and introduce the multiscale template descriptors with an efficient matching method using feature distribution while learning. Next, we focus on target objects with edge structures, and describe edge-based feature descriptors and demonstrate their use on real-time augmented reality applications under viewpoint changes and occlusions. We conclude this thesis with discussions on extension to multi-modal representations and relations to visual exploration tasks.