Stream-weighted HMM for audio-visual ASR: a study on connected digit recognition
Michael T. Chan · 2004
We present some new results on connected digit recognition in noisy environments by audio-visual speech recognition. We derive hybrid (geometric- and appearance-based) visual lip features using a real-time lip-tracking algorithm that we proposed previously. Using a single-speaker corpus modeled after the TIDIGITS database, we build whole-word HMMs using both single-stream and 2-stream modeling strategies. For the 2-stream HMM method, we use stream dependent weights to adjust the relative contributions of the two feature streams based on the acoustic SNR level. The 2-stream HMM consistently gave the lowest WER, with an error reduction of 83% at -3 dB SNR level compared to the acoustic-only baseline. Visual-only ASR WER at 6.85% was also achieved, showing the effectiveness of the visual features. A real-time system prototype was developed for concept demonstration.