Investigation of Automatic Video Summarization using Viewer’s Physiological, Facial and Attentional Features
Sergio Cavalcanti de Paiva, Herman Martins Gomes · 2019
Video summarization aims at the selection of a concise and representative set of keyframes or video segments that allows the identification of the video content. In either cases, traditional summarization techniques usually work by segmenting the video into shots, representing video frames as feature vectors of color, texture, audio, among other features, clustering frames with similar features and selecting most representative keyframes or segments, sometimes guided by a video-to-summary ratio target. The resulting summaries are typically subject independent and do not take into account specific viewer's behavior. Instead of using intrinsic features extracted from the video for summarization, in this article we study whether personalized (subject dependent) video summaries can be obtained from physiological, facial, and attentional data captured from the viewers. More specifically, we study the relationship between personalized video summaries reported by viewers and their data captured during the display of different video genres. A dataset of fifteen videos was used in the experiments. During the exhibition of the videos, the viewer's physiological, facial, and attentional data were recorded, analyzed and synchronized. Several machine learning models were trained to test our hypothesis. We obtained k-fold cross validation accuracies that were above the chance for the best learned models. As a result of this study, we conclude that it is possible to train a learning machine that can produce customized summaries that are closer to user preferences compared to randomly produced summaries.