Confusability of Phonemes Grouped According to their Viseme Classes in Noisy Environments
Patrick J. Lucey, Terry Z. Martin, Subramanian Sridharan · QUT ePrints (Queensland University of Technology) · 2004
Using visual information, such as lip shapes and movements, as the secondary source of speech information has been shown to make speech recognition systems more robust to problems associated with environmental noise, training/testing mismatch and channel and speech style variations. Research into utilising visual information for speech recognition has been ongoing for 20 years, however over this period, a study into which visual information is the most useful or pertinent to improving speech recognition systems has yet to be performed. This paper presents a study to determine the confusability of the phonemes grouped into their viseme classes over various levels of noise in the audio domain. The rationale behind this approach is that by establishing the interclass confusion for a group of phonemes in their viseme class, a better understanding can be obtained on the complementary nature of the separate audio and visual information sources and this can be subsequently applied in the fusion stage of an audio-visual speech processing (AVSP) system. The experiments performed show high interclass confusion variability at the 0dB and-6dB SNR levels. Further analysis found that this was mainly due to a phonetic imbalance in the dataset. Due to this result, it was suggested that it would be appropriate for an AVSP system used for digit recognition applications heavily weight the visual modality for the phonemes that are most prevalent such as the phoneme N. 1.