HMM-based audio-visual speech recognition integrating geometric- and appearance-based visual features
Michael T. Chan · 2002
A good front end for visual feature extraction is an important element of audio-visual speech recognition systems. We propose a new visual feature representation that combines both geometric- and pixel-based features. Using our previously developed contour-based lip-tracking algorithm, geometric features including the height and width of the lips are automatically extracted. Lip boundary tracking allows accurate determination of a region of interest from which we construct pixel-based features that are robust to variation in scale and translation. Motivated by computational considerations, we selected a subset of the pixels in the center of the inner mouth area that was found to capture sufficient details of the appearance of the teeth and tongue for assisting in the discrimination of spoken words. We show the advantage of the combination of these visual features for visual-only and audio-visual speech recognition of isolated digits.