DCT-based video features for audio-visual speech recognition

Martin Heckmann, Kristian Kroschel, Christophe Savariaux, Frédéric Berthommier · 2002

Encouraged by the good performance of the DCT in audio-visual speech recognition [1], we investigate how the selec-tion of the DCT features influences the recognition scores in a hybrid ANN/HMM audio-visual speech recognition sys-tem on a continuous word recognition task with a vocab-ulary of 30 numbers. Three sets of features, based on the mean energy, the variance and the variance relative to the mean value, were chosen. The performance of these fea-tures is evaluated in a video only and an audio-visual recog-nition scenario with varying Signal to Noise Ratios (SNR). The audio-visual tests are performed with 5 types of addi-tional noise at 12 SNR values each. Furthermore the results of the DCT based recognition are compared to those ob-tained via chroma-keyed geometric lip features [2]. In order to achieve this comparison, a second audio-visual database without chroma-key has been recorded. This database has similar content but a different speaker. 1.

Read the paper · More papers on PaperTik