Normalized training for HMM-Based visual speech recognition
Yoshihiko Nankaku, Keiichi Tokuda, Tadashi Kitamura, Takao Kobayashi · Electronics and Communications in Japan (Part III Fundamental Electronic Science) · 2006
This paper discusses parameter estimation for a continuous density HMM (hidden Markov model) method of visual speech recognition. Past studies of visual speech recognition can be broadly divided into two approaches, the image-based method and the model-based method. The image-based method is a method in which some preprocessing, such as subsampling and principal component analysis, is applied to the pixel values of the original image, and the result is used as the feature vector. In this approach, the position and size of the lips and the illumination conditions have a direct effect on the recognition rate. Thus, the normalization of these factors is the basic technique. The ordinary approach in the conventional normalization is to provide a criterion independently of the HMM, and to apply normalization before learning. In this paper, normalization by the ML (maximum likelihood) criterion is considered. Normalized training is proposed in which the normalization processes for elements such as the position, size, inclination, mean brightness, and contrast of the lips are integrated with the training of the model. The proposed method is formulated on the basis of an EM (expectation maximization) algorithm in which monotonically increasing behavior of the likelihood of the training data is guaranteed, by iteration of normalized training. The effectiveness of the proposed method is demonstrated in a word recognition experiment using the M2VTS database. © 2006 Wiley Periodicals, Inc. Electron Comm Jpn Pt 3, 89(11): 40–50, 2006; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/ecjc.20281