Towards Cross-Media Feature Extraction.

Thierry Declerck, Paul Buitelaar, Jan Nemrava, David A. Sadlier · 2008

In this paper we describe past and present work dealing with the use of textual resources, out of which semantic information can be extracted in order to provide for seman-tic annotation and indexing of associated image or video material. Since the emergence of semantic web technolo-gies and resources, entities, relations and events extracted from textual resources by means of Information Extraction (IE) can now be marked up with semantic classes derived from ontologies, and those classes can be used for the se-mantic annotation and indexing of related image and video material. More recently our work aims additionally at tak-ing into account extracted Audio-Video (A/V) features (such as motion, audio-pitch, close-up, etc.) to be com-bined with the results of Ontology-Based Information Ex-traction for the annotation and indexing of specific event types. As extraction of A/V features is then supported by textual evidence, and possibly also the other way around, our work can be considered as going towards a “cross-media feature extraction”, which can be guided by shared ontologies (Multimedia, Linguistic and Domain ontolo-gies).

Read the paper · More papers on PaperTik