Multimedia knowledge: discovery, classification, browsing, and retrieval
Ana Belen Benitez Jimenez, Shih‐Fu Chang · 2005
Humans behave and reason by learning and maintaining dynamic models of the world. In addition, there is evidence from psychology that human models contain textual nodes (semantic; word “watermelon”) and audio-visual nodes (perceptual; quality of being green). This is why in an attempt to make sense of multimedia as humans do, this thesis focuses on representing and discovering semantic and perceptual knowledge about the world using multimedia and knowledge from external resources such as the electronic dictionary WordNet. This thesis also proposes techniques that use the discovered knowledge to advance multimedia classification, browsing, and retrieval applications. In this thesis, we propose a unified framework, MediaNet, that uses multimedia for representing semantic and perceptual knowledge in the form of concepts networks with media examples, which we call “medianets”. The MediaNet framework extends semantic knowledge frameworks such as thesaurus and ontologies by including perceptual knowledge, and exemplifying concepts and relationships using multimedia. We use the MPEG-7 standard to represent medianets in an interoperable way. This thesis also proposes new techniques for discovering and summarizing medianets from annotated images. Medianets are constructed by clustering images based on visual and textual features; and disambiguating the senses of words in annotations using WordNet and the image clusters. Visual, statistical, and semantic relationships are then discovered between clusters and senses (i.e., concepts). Finally, medianets can be summarized by merging statistically similar concepts. In contrast to prior work, we integrate the processing of images and annotations to improve the concept discovery In addition, our summarization techniques are general, automatic, and applicable to any domain. Image classifiers can be used to annotate images with semantic labels such as “person”, “mountain”, and “outdoors”. A key problem of current image classification systems is the lack of systematic methods for finding relevant classes and their relations. We propose the use of extracted medianets for automatically discovering salient classes and combining several individual detectors for superior accuracy (up to a 15% gain). We train individual detectors to predict the presence of concepts in images, which are combined statistically using a Bayesian network based on the medianets. Image browsers enable users to gain a quick insight into the content of a collection by supporting several exploration tasks (e.g., locating images). Current approaches organize images on large and complex structures, which result in inefficient navigation. (Abstract shortened by UMI.)