Location normalization of HMM-based lip-reading: experiments for the M2VTS database
Oscar Vanegas, Keiichi Tokuda, Tadashi Kitamura · 1999
This paper describes an HMM-based lip location normalization process, in order to improve the recognition performance in automatic lip-reading. This paper uses the image-based method in order to represent the lip visual information. One of the most critical factors which affect the recognition results in image-based method is the position of lips in frames. This paper describes a method to normalize the lip location which is similar to SAT (speaker adaptive training), and presents several experiments which were carried out in order to measure the effectiveness of the proposed method. Experiments of isolated words with and without the original movement from speakers were carried out on the M2VTS database.