Audio-Visual Voice Activity Detection Based on an Utterance State Transition Model
Takami Yoshida, Kazuhiro Nakadai · Advanced Robotics · 2012
This paper describes improvements in Audio-Visual Voice Activity Detection (AV-VAD) using a state transition model. Audio-Visual integration is a promising approach to improve the noise-robustness of VAD. Our proposed AV-VAD is based on a state transition model with four utterance states: speech, nonspeech, and two asynchronous states, a beginning motion state and an ending motion state, corresponding to lip activity before and after voice activity, respectively. We implanted a prototype system into an upper torso humanoid SIG and evaluated the ability of the proposed method by using auditory- and/or visually-contaminated data. Experimental results showed that the robustness of the VAD improved even when the resolution of images was low. On average, our proposed method resulted in a 7.8 point improvement in word detection rates compared with existing AV-VAD methods.