Viseme recognition - a comparative study

Mariusz Leszczyński, Władysław Skarbek · 2006

Three classification algorithms for visual mouth appearances (visemes) which correspond to phonemes and their speech contexts, were compared w.rt. recognition rate, time complexity, and ROC performance. Two feature extraction procedures were verified. The first one is based on the normalized triangle MESH covering mouth area and the color image texture vector indexed by barycentric coordinates. The second procedure performs DFT on the image rectangle including mouth w.rt. small blocks of DFT coefficients. The classifiers has been designed by PCA approach and by the optimized LDA method which uses two singular subspaces approach. It appears that DFT+LDA exhibits higher recognition rate than MESH+LDA and MESH+PCA methods - 97.6% versus 94.4 and 90.2%, respectively. It is also much faster than MESH+PCA (5 ms per one video frame versus 26 ms on Pentium IV, 3.2 GHz) and slower than MESH+LDA (5 ms versus 1 ms).

Read the paper · More papers on PaperTik