Talking face: using facial feature detection and image transformations for visual speech

A. Arya, Babak Hamidzadeh · 2002

Visual presentation of a talking person requires the generation of image frames showing the speaker in various views while pronouncing various phonemes. The existing approaches, mostly use either a complex 3D geometric model to reconstruct a desired image or a set of 2D images for each viewpoint, to select from. We propose a new system which utilizes facial feature detection and image-based transformation to create any talking frame using only one given image from the desired viewpoint and a set of reference images from one standard view. The proposed approach, together with optical flow-based view morphing and a customizable concatenative text-to-speech, makes a personalized visual speech generation system which can be used for moving/talking head applications where an optimal trade-of between computational complexity and image database requirements is necessary.

Read the paper · More papers on PaperTik