Detection of documentary scene changes by audio-visual fusion
Atulya Velivelli, Chong‐Wah Ngo, Thomas S. Huang · 2003
Abstract. The concept of a documentary scene was inferred from the audio-visual characteristics of certain documentary videos. It was ob-served that the amount of information from the visual component alone was not enough to convey a semantic context to most portions of these videos, but a joint observation of the visual component and the audio component conveyed a better semantic context. From the observations that we made on the video data, we generated an audio score and a vi-sual score. We later generated a weighted audio-visual score within an interval and adaptively expanded or shrunk this interval until we found a local maximum score value. The video ultimately will be divided into a set of intervals that correspond to the documentary scenes in the video. After we obtained a set of documentary scenes, we made a check for any redundant detections. 1