Video Editing Based on Situation Awareness from Voice Information and Face Emotion

Tetsuya Takiguchi, Jun Adachi, Yasuo Ariki · InTech eBooks · 2010

In this chapter, we investigated about home video editing based on audio with a twochannel (stereo) microphone and facial expression, where the video content is automatically recorded without a cameraman. In order to capture a talking person only, a novel voice/non-voice detection algorithm using AdaBoost, which can achieve extremely high detection rates in noisy environments, is used. In addition, the sound source direction is estimated by the CSP (Crosspower-Spectrum Phase) method in order to zoom in on the talking person by clipping frames from videos, where a two-channel (stereo) microphone is used to obtain information about time differences between the microphones. Also, we extract facial feature points by EBGM (Elastic Bunch Graph Matching) to estimate atmosphere class by SVM (Support Vector Machine). When the atmosphere of the other person except the speaker is not "positive" class, the digital camera work zooms in only the speaker. When the atmosphere of the other person is "positive" class, a wide shot is taken. Our proposed system can not only produce the video content but also retrieve the scene in the video content by utilizing the detected voice interval or information of a talking person as indices. To make the system more advanced, we will develop the sound source estimation and emotion recognition in future, and we will evaluate the proposed method on more test data.

Read the paper · More papers on PaperTik