Extracting visual information from text: using captions to label faces in newspaper photographs
Rohini K. Srihari · 1992
There are many situations where linguistic and pictorial data are jointly presented to communicate information. In the general case, each of these two sources conveys orthogonal information. A computer model for integrating information from the two sources requires an initial interpretation of both the text and the picture followed by consolidation of information. The problem of performing general-purpose vision without apriori knowledge (needed in such a situation) is nearly impossible. However, in some situations, the text describes salient aspects of the picture. In such situations, it is possible to extract visual information from the text, resulting in a conceptualised graph describing the structure of the accompanying picture. This graph can then be used by a computer vision system in the top-down interpretation of the picture. In this dissertation, a computational model for understanding pictures based on information in accompanying captions is presented. The use of SNePS (Semantic Network Processing System) as the common intermediate representation for both linguistic and pictorial information is discussed. Specifically, we present the knowledge representations and interpretations that comprise the model. The focus of this dissertation is on the generation of a conceptualised graph, a SNePS network which reflects a cognitive agent's conceptualisation of a picture based on information contained in a descriptive caption. This representation includes information about objects appearing in the picture and spatial constraints between them, information used in the subsequent task of labelling objects in the picture. A substantial portion of the dissertation is devoted to techniques of extracting such visual information from text. The techniques are based on both syntactic and semantic considerations. We classify linguistic methods of identifying objects in pictures into several broad categories and, for each category, discuss the manner in which visual information can be extracted. The problem of dynamically generating model descriptions for objects (and for entire pictures) is illustrated. A theoretical solution to this problem is presented and illustrated through an example. As a test of the model, we present an implementation, PICTION, whereby information obtained from parsing a caption of a newspaper photograph is used to identify human faces in the photograph. A key component of the system is the utilisation of spatial constraints in order to reduce the number of possible labels that could be associated with a face. These constraints are generated by a natural-language processing system that examines the caption in detail. We report on the extensive testing of the system and discuss the results obtained. The method of evaluating the performance of PICTION can be used by any face-identification system.