Qualitative scene descriptions from images for integrated speech and image understanding
Gudrun Socher · PUB – Publications at Bielefeld University (Bielefeld University) · 1997
Human-computer interaction using means of communication which are natural to humans, like spoken instructions or gestures, has always been a challenging task. In this thesis, we address the subproblem of fusing the understanding of spoken instructions with the visual perception of the environment. We describe the design and implementation of a high-level computer vision component for the integrated speech and image understanding system QUASI-ACE. QUASI-ACE is a prototype of a ‘situated artificial communicator’, a system which aims to interact with humans in a natural way given a specific scenario or situation. A toy assembly scenario is our domain. The system QUASI-ACE is able to identify objects intended in spoken instructions given by a human instructor, based on results from the image understanding component which visually observes the scene. The high-level image understanding is accomplished by first reconstructing the 3D scene from uncalibrated stereo images. We use a model-based approach fitting 3D object models to 2D features extracted from images to estimate the pose of all objects and the camera parameters. In this context we developed a new method of using ellipses for 3D reconstruction based on projective invariants. The second image understanding step involves the computation of the qualitative features ‘type’, ‘color’, and ‘spatial relations’ from the 3D data and from 2D object hypotheses obtained from object recognition. These qualitative features are represented as fuzzified vectors assigning a likelihood value to each category in the feature space. The identification of the intended objects in the spoken instructions is based on a Bayesian network approach. The objects with the highest joint probability of being observed in the scene and being intended in the instructions are identified using the common qualitative representation for observed ‘type’, ‘color’, and ‘spatial relations’ as well as the uttered ‘type’, ‘color’, ‘size’, ‘shape’, and ‘spatial relations’. Domain and prior knowledge is incorporated in the identification process. We show the performance and robustness of our image understanding and object identification modules in the context of the entire system QUASI-ACE using real images and spoken instructions.