From speech to talking faces: lip movements estimation based on linear approximators

Fabrizio Vignoli · 2002

In human communication, speech understanding is greatly improved by the bimodal acoustic-visual effect, with respect to simple speech. This is particularly clear when the communication takes place in noisy environments or for non-native speakers. In this paper, we propose a novel algorithm based on linear approximators that estimates the lip movements from a timed sequence of phonemes. This sequence can be generated from real speech, by a segmentation technique based on a hidden Markov model (HMM), or from a text-to-speech system. The algorithm consists of two modules: the training module and the synthesis module. The training module is based on a eigen-analysis of an audiovisual database recorded for this purpose. The synthesis module takes as input the sequence of phonemes and implements an implicit coarticulation model. A later post-processing step converts the parameters estimated into a sequence of facial animation parameters that are compliant to the new MPEG-4 standard. The algorithm has been tested with FAE (Facial Animation Engine), which is an MPEG-4 compliant system developed at the author's university.

Read the paper · More papers on PaperTik