Integrating Multimedia Information Retrieval, Signal Processing and Human Computer Interaction
Carlos Teixeira · Portuguese National Funding Agency for Science, Research and Technology (RCAAP Project by FCT) · 2008
This abstract presents current research goals of the author, most of which were built within the framework of the author's research group HCIM (Human-Computer Interaction and Multimedia Research Team at LASIGE-DI/FCUL). This group had long term projects in the area of digital talking books and hypermedia (Chambel and Guimaraes, 2002). This was an inspiration for the creation of new multimedia contents and tools intertwining audio-visual and textual related material, linking different expression modes coherently in a “multimedia object”. Adding higher level information that cuts across the several modalities (text, audio, images and film) significantly augment the range of functionalities to be considered. Higher levels of information will then be used to provide high level summaries or multimedia indexes of the (mixed) content, drawing from work in ontology extraction from text, text classification and natural language generation techniques (Lawrie and Croft, 2003). Using relationships between text (near or aligned) and pictures or video-clip can increase connectivity between the different modalities and provide truly multimedia abstracts. Research is conducted on the subjects of auto-illustrate (link images or videos to text) and auto-annotate (link text to audiovisual objects, Barnard et al. 2003). Drawing from work in areas as diverse as computational stylistics, ontology validation and terminological systems, it is now possible to assign a set of labels to a text, covering widely different dimensions: • text structure (introduction, action, commentary, dialogue, different scenes in a fiction book, or in a carefully edited TV program, etc.) • events (roles, date, place, causes and consequences) • topic, subject, news cluster, named entities, geographical scope, etc., and • establishing links between subjects, events and texts The goal here is to go one step further than systems which allow people to access individual scenes, such as MUSCLE’s example of “give me all video clips where JK is to the right of his assistant” (Enis Cetin, 2005) to more abstract (and arguably more useful) notions such as “gun firings in Iraq” (news as raw material) or “TV debates where the consequences of the oil raising price are named” or “pictures of Algarve this Summer” for other kinds of plausibly r elevant information access for different kinds of users. Several applications are envisaged for providing integrated multimedia browsing, querying and production of the new contents. In a first phase, partially annotated related multimedia contents are selected in order to allow the construction of new integrated contents. This requires the development of specialized text aligning algorithms. A new text alignment was presented in (Teixeira and Respicio, 2007), not based on the classical longest subsequence approach, with new features supporting new applications. Further advanced techniques such as automated knowledge discovery and extraction, topic detection and automatic summarization, are expected to improve this applications.