Audio-visual continuous speech recognition using a coupled hidden Markov model
Xiaoxing Liu, Yibao Zhao, Xiaobo Pi, Luhong Liang, Ara Nefian · 2002
With the increase in the computational complexity of recent computers, audio-visual speech recognition (AVSR) became an attractive research topic that can lead to a robust solution for speech recognition in noisy environments. In the audio visual continuous speech recognition system presented in this paper, the audio and visual observation sequences are integrated using a coupled hidden Markov model (CHMM). The statistical properties of the CHMM can describe the asyncrony of the audio and visual features while preserv-ing their natural correlation over time. The experimental re-sults show that the current system tested on the XM2VTS database reduces the error rate of the audio only speech recognition system at SNR of 0db by over 55%. 1.