Video summaries and cross-referencing

John R. Kender, A. Aner · 2002

A system is presented to construct a highly compact hierarchical representation of video. It is based on a tree-like representation where the bottom level is composed of frames, and the highest level represents a non-temporal segmentation of the video, and is suitable for well-structured video genres. We first propose a method to segment the video into shots, which is based on a leaky memory model, previously proposed for scene transition detection. It is unique since it presents a unified approach for shot and scene detection, and for key-frame selection. We next demonstrate the benefits of using mosaics for representing shots. A novel method for mosaic alignment and comparison is proposed, which is shown to be both efficient and effective. A scene distance measure based on mosaic comparison is then defined and used to cluster the scenes into a higher abstraction of video content, the physical settings. We demonstrate this hierarchical representation using situation comedies, in which this abstraction has a strong semantic meaning in summarizing videos. By comparing physical settings across different episodes of the same situation comedy, we determine the main plots of each episode. We conclude by demonstrating a video summary tool specialized for fast browsing of situation comedies. In another example, we apply our mosaic comparison method to detect significant events in basketball games, enabling fast forwarding from one event to the next.

Read the paper · More papers on PaperTik