Multi-modal Video Summarization

Jia-Hong Huang · 2024

The increasing prevalence of video content on platforms like YouTube and Vimeo has led to a growing demand for effective video summarization techniques. These methods aim to extract relevant content from videos and present it in a condensed form, addressing the challenge of information overload faced by users. However, the quality of the generated summaries is crucial, especially in fields like forensics, journalism, and sports analysis, where accurate and concise summaries are essential for decision-making and analysis. Traditional video summarization methods often rely solely on visual information, limiting their ability to capture textual cues and emotional content present in videos. To overcome this limitation, this thesis proposes leveraging text-based queries as context to enhance the effectiveness of video summarization. By incorporating textual information alongside visual data, the proposed approach aims to generate query-dependent video summaries that better align with users' requirements. Furthermore, the thesis leverages a novel conditional modeling perspective to impart a more human-like quality to the video summarization process. These methods hold promise for various applications, including documentary filmmaking and educational content creation. Additionally, the thesis addresses the challenge of data scarcity in video summarization by proposing a self-supervised learning approach that leverages pretext tasks to generate pseudo-labels for model training. In 2020, we proposed a query-controllable video summarization technology, which is elaborated upon in the second chapter of this thesis. Notably, Google has adopted a similar feature, integrating it into their large language model (LLM) Gemini Pro-1.5 version, in 2024.

Read the paper · More papers on PaperTik