Indoor scene understanding

Chandra Kambhamettu, Gowri Somanath · 2012

Intelligent robots that interact and blend in with humans have been one of the core goals of AI. One of the important sensors on such a robot is its vision system. Humans are known to derive maximal sensory feedback and experience from what they see. To enable a capability to imitate a fraction of the interaction and derived knowledge is a challenging task. It has become clear that the task involves devising algorithms that allow for adaptive and incremental instance based learning. We study some of the basic problems in designing such vision systems, specifically for indoor environments. Specifically, this thesis seeks to address (1) object discovery through unsupervised scene segmentation, (2) organization of object library for efficient recognition and retrieval, (3) purpose based object classification and part labeling, and (4) scene classification and place recognition. We present novel approaches and algorithms towards solutions for the above problems. We also apply some of the techniques to solve problems in facial analysis and outdoor scenes, thus demonstrating the generality of the proposed schemes. The key motivation in our work has been the use of geometry at various levels and the close relation to findings from various cognitive science studies. We have shown that knowledge of scene and object geometry can significantly improve performance when compared to previous state-of-the-art methods. Our experiments show that use of depth, 3D structure and shape can prove robust and effective for practical solutions. We have used geometry and properties derived thereof at various degrees of specificity and abstraction. Further we generalize the use of geometry to include notion of arrangement in scenes. With advent of better computing and imaging resources, using cameras on robotic systems is becoming increasingly popular. We have designed a cost effective lens adapter which can be used to convert any single camera to a stereo camera, to handle cases where multiple camera cannot be mounted. Thus geometry can be effectively acquired to allow the use of our approaches. In addition to prior work in scene understanding area of computer vision, many of our algorithms have also been inspired from theories and findings from human vision and perception studies. We demonstrate that practical and computationally efficient machine vision systems can be based on concepts from cognitive science. We have taken inspiration from seminal works of Palmer, Rosch, Minsky, Tversky and many others. We observed that many results from our and related works correlate with hypothesis and findings from human vision. Though the underlying representations and mechanisms are different, at a conceptual level we may replicate certain processes of human vision into machines, thus contributing to the success of many vision algorithms including ours.

Read the paper · More papers on PaperTik