On audio-visual synchronization for viseme-based speech synthesis

Jianxia Xue, Abeer A. Alwan, Edward T. Auer, Lynne E. Bernstein · The Journal of the Acoustical Society of America · 2004

In viseme-based visual speech synthesis [e.g., T. Ezzat and T. Poggio, Int. J. Comput. Vision 38, 45–57 (2000)], it is frequently assumed that each viseme can be represented by an image which corresponds to the temporal mid-point of the phoneme’s acoustic segment. We investigated the validity of this assumption in a study of 3-D motion data extracted from marker positions on the face of one talker. Using principle components analysis (PCA), the first five principle motion tracks comprising the lips and the jaw were studied. The temporal positions of the local extrema in the principle motion tracks were analyzed. Statistics of the distances between these local extrema and the mid-points of the corresponding acoustic segments were compared. Results showed that the local extrema of the motion tracks were, for the most part, not well-aligned to the phoneme’s acoustic midpoint with a few exceptions. For example, for /s/ the local extrema were well aligned (small variance), while for /m/ and /f/, alignments pattern could be found for several, but not all, tokens. The results suggest that phonetic context and speaking rate must be taken into account when characterizing facial configurations in visual speech synthesis. [Work supported in part by the NSF.]

Read the paper · More papers on PaperTik