Audio-Visual and EEG-Based Attention Modeling for Extraction of Affective Video Content
Irfan Mehmood, Muhammad Husain As Sajjad, Sung Wook Baik, Seungmin Rho · 2015
Video summarization is a procedure to reduce redundancy and generate concise representation of the video data. Extracting affective key frames from video sequences is an enthusiastic approach amongst video summarization schemes. Affective key frames refer to the intensity and type of feelings that are contained in video and expected to arise in spectators mind. Recent summarization schemes consider audio and visual information. However, these data modalities are not sufficient to accurately perceive human attention, failing to extract semantically relevant content. Video content incites strong neural responses in users, which can be measured by analyzing electroencephalography (EEG) brain signals. Merging EEG and multimedia analysis can serve as a bridge, linking the digital representation of multimedia and user perception. In this context, we propose an affective video content extraction scheme that integrates human neuronal signals with audio-visual features for better perception and comprehension of digital videos. Experimental results shown that the propose model can accurately reflect user preferences, and facilitate extraction of highly affective and personalized summaries.