Multimodal Fusion Strategies

Hanseok Ko · 2018

Two-hour movie or a short movie clip as its subset is intended to capture and present a meaningful (or significant) story in video to be recognized and understood by human audience. What if we substitute the task of human audience with that of an intelligent machine or robot capable of capturing and processing the semantic information in terms of audio and video cues contained in the video? By using both auditory and visual means, human brain processes the audio (sound, speech) and video (background image scene, moving video objects, written characters) modalities to extract the spatial and temporal semantic information, that are contextually complementary and robust. Smart machines equipped with audiovisual multisensors (e.g. CCTV equipped with cameras and microphones) should be capable of achieving the same task. An appropriate fusion strategy combining the audio and visual information would be a key component in developing such artificial general intelligent (AGI) systems. This talk reviews the challenges of current video analytics schemes and explores various sensor fusion techniques [1, 2, 3, 4] to combine the audio-visual information cues for video content analytics task.

Read the paper · More papers on PaperTik