A Conceptual Framework for High-level Vision FBI-HH-B-245/03

Fachbereich Informatik · 2003

Abstract In this report we present essential elements of a conceptual framework for high-level vision(HLV). The scope of HLV is defined as scene interpretation above the level of objectrecognition. It is shown that models, on which such interpretations can be based, typicallydescribe aggregates composed of meaningful parts, related to each other by temporal andspatial constraints. A frame-based representation is proposed which is based on techniquesimported from configuration methodology. The hypothesise-and-test interpretation process isdescribed for a table-laying example. It is shown that expectations are generated by part-whole reasoning. A temporal constraint net is proposed for the incremental evaluation ofqualitative temporal constraints. A similar approach is sketched for spatial constraints whichare represented by grid locations in a reference frame attached to an object. Scope of high-level vision (HLV) In this section we review developments in Computer Vision which contribute to a widerunderstanding of the vision task as compared to classical vision tasks such as recognizing ortracking single objects. These developments have to be taken into consideration whendesigning the conceptual framework for HLV in a cognitive agent as envisioned in the projectCogVis.From human vision it is evident that what we see is interpreted in the light of diverseknowledge and of experiences about the world. The scope of this knowledge - often termedcommon-sense knowledge - can best be seen when we consider silent-movie watching as aComputer Vision task, for example, watching and understanding a film with Buster Keaton. Ifa vision system were to interpret the visual information of such a film in a depth comparableto humans, the system would have to resort to knowledge about typical (and atypical)behaviour of people, intentions and desires, events which may happen, everyday physics, thenecessities of daily life etc. This is knowledge far beyond the visual appearance of singleobjects, and a vision system capable of silent-movie understanding clearly has to solve tasksbeyond single-object recognition.As early as 1955 Computer Vision has been proposed as a task integrated in a cognitivecontext [Selfridge 55] and interacting with other cognitive processes. But Computer Visionresearch was in its infancy then, and a much narrower view of the vision task had to bepursued for several decades. The idea of integrating vision with other cognitive processes wasactively investigated for the first time in the eighties in projects dealing with natural-languagedescriptions of imagery [vHahn et al. 80, Nagel 88, Neumann 89]. One of the importantinsights of this work was that qualitative descriptions had to be derived from geometric scenedescriptions as an interface to language and symbolic reasoning.

Read the paper · More papers on PaperTik