ADAPTIVE BIMODAL SENSOR FUSION FOR AUTOI/IATIC SPEECHREADING

Uwe Meiel · 2004

2. SYSTEM DESCRIPTION We present recent work on improving the performance of automated speech recognizers by using additional visual information (Lip-/Speechreading), achieving error reduction of up to 50%. This paper focuses on different methods of combining the visual and acoustic data to improve the recognition performance. We show this on an extension of an existing state-of-the-art speech recognition system, a modular MS-TDNN. We have developed adaptive conibination methods at several levels of the recognition network. Additional information such as estimated signal-to-noise ratio (SNR) is used in some cases. The results of the different conibination methods are shown for clean speech and data with artificial noise (white, music, motor). The new combination methods adapt automatically to varying noise In the basic set-up, we record, in parallel, the acoustic speech and the corresponding serieij of mouth images of the speaker. The speaker and his lip!; are found and tracked automatically. We use speaker-dependent continuous spelling of German letter strings (26 letter alphabet) as our task. Words in our database are 8 letters long on average. I noise signal-to-noise ratio I clean motor 25 dB and 10 dB Table 1. Acoustic environments tested (dB SNR). conditions making hand-tuned parameters unnecessary.

Read the paper · More papers on PaperTik