Semantic mosaic for indexing and compressing instructional videos
Tianhui Liu, John R. Kender · 2004
A new approach for content analysis and semantic compression of instructional videos is presented. Cameras often only capture a portion of a hand-drawn slide or blackboard panel that dominate this genre. We therefore have designed a novel semantic technique that retrieves the teaching content (text lines, figures, emphasis marks, etc.) visible in a frame, de-skews them and enhances their contrast, and stitches them back together into "virtual" slides that maintain the relative spatial relations of the captured fragments. Since this semantic mosaicing does not need precise pixel-wise matching, it has low computational complexity. These virtual slides maintain the useful teaching content of their real-world counterparts, and we further stitch them together to form a content summary slide to summarizes an extended video segment. By detecting content and recording its temporal development, we can then semantically compress a video into several such higher quality mosaics, and reconstruct the original instructional video by displaying the mosaics instead at the appropriate times. Our design runs in one pass and in real time. We apply the system to a 17-minute instructional video, and show the results of its semantic compression into virtual slide mosaics.