Audio-visual automatic speech recognition using Dynamic Bayesian Networks
Helge Reikeras · SUNScholar (Stellenbosch University) · 2011
Abstract—In audio-visual automatic speech recognition (AVASR) both acoustic and visual modalities of speech are used to determine what a speaker is saying. In this paper we propose a basic AVASR system that uses mel-frequency cepstrum coefficients (MFCCs) as acoustic features, active appearance model (AAM) parameters as visual features, and dynamic Bayesian Networks (DBNs) as probabilistic models of audio-visual speech. The performance of the AVASR system is tested using the Clemson University audio-visual experiments (CUAVE) database. We find that, as expected, visual-only speech recognition (au-tomatic lip-reading) performs worse than audio-only speech recognition. However, by integrating visual and acoustic speech information we are able to significantly increase performance, in particular in noisy acoustic environments. I.