DETECTION OF SPEECH EVENTS IN REAL ENVIRONMENTS THROUGH FUSION OF AUDIO AND VIDEO INFORMATION USING BAYESIAN NETWORKS

Takashi Yoshimura, Futoshi Asano, Youichi MOTOMURA, Hideki Asoh, Naoyuki Ichimura, Kiyoshi Yamamoto, Satoshi Nakamura · 2003

A method of combining audio and video information for detecting and separating speech events in real environments is presented. The method is effective for automatic speech recognition under conditions of multiple sound sources. Sound localization using a microphone array and human tracking by stereo vision are combined using a Bayesian network for detecting speech events. Based on the detected information, the time and location of speech events, a maximum likelihood adaptive beamformer is constructed and the speech signal is separated from background noise and interference. An advantage of using the Bayesian network is that the scheme allows the correspondence of audio and video coordinates to be established with ambiguity by modeling a joint probability distribution. The results of off-line experiments in a real environment with television and/or music interference are presented as verification of the proposed scheme. 1.

Read the paper · More papers on PaperTik