From mid-level to high-level: Semantic inference for multimedia retrieval

Qianni Zhang, Ebroul Izquierdo · 2010

The problem of bridging the semantic gap can be approached by dividing all types of metadata extracted from multimedia content into three levels - low, mid and high - according to their levels of semantic abstraction and try to define the mapping between them. This paper proposes a scheme for extracting high-level semantic information out of mid-level features, which can be applied in dealing with highly semantic queries in image retrieval. Mid-level features used in this research contain some level of semantic meaning but are not directly useful in real retrieval scenarios. However, they usually have strong relationships to high-level queries but these relationships are often ignored due to their implicitness. The aim of the proposed approach is to explore hidden interrelationships between mid-level features and the high-level query terms, by learning a Bayesian network model from a small amount of training data. Semantic inference and reasoning is then carried out based on the learned Bayesian network model, in order to decide whether a video is relevant to a high-level query. The extracted high-level semantic terms can be annotated on the video content for future retrieval. Two experimental scenarios were considered in this paper and the experiments on RUSHES videos have produced satisfactory results.

Read the paper · More papers on PaperTik