Retrieving Video Segments Based on Combined Text, Speech and Image Processing

Harris Papageorgiou, Athanassios Protopapas, Thomas Netousek · 2003

This paper describes a multimedia, multilingual and multimodal research system (CIMWOS) supporting content-based indexing, archiving, retrieval and ondemand delivery of audiovisual content. There are several projects, aiming at developing advanced technologies and systems to tackle the problems encountered in multimedia archiving and indexing [8], [9], [10]. CIMWOS [1] (Combined IMage and WOrd Spotting) incorporates an extensive set of multimedia technologies by seamless integration of three major subsystems – text, speech and image processing – producing a rich collection of XML metadata annotations following the MPEG-7 standard. These XML annotations are further merged and loaded into the CIMWOS Multimedia Database. Additionally, they can be dynamically transformed for interchanging semantic-based information into RDF and Topic Maps documents via XSL stylesheets. The CIMWOS Retrieval Engine is based on a weighted boolean model with intelligent indexing components. An ergonomic and userfriendly web-based interface allows users to efficiently retrieve video segments by a combination of media description, content metadata and natural language text. The database is a large collection of broadcast news and documentaries in three languages (English, Greek, and French), while the open architecture allows for more languages to be incorporated in the future.

Read the paper · More papers on PaperTik