Speech Driven Synthesis of Talking Head Sequences
Peter Eisert, Subhasis Chaudhuri, Bernd Girod · 1997
3D Image Analysis and Synthesis, pp. 51-56, Erlangen, November 1997. In this paper we present a method for generating talking head sequences directly from speech signals. A neural network is used to derive MPEG4 facial animation parameters related to a person 's mouth shape with low computational complexity. For the training of the network we use estimated parameters from an iterative and linear algorithm that uses 2D point correspondences of marker positions in video sequences. Having estimated the facial parameters we can render an animation of a speaking arbitrary person. Experimental results show that the appearance of the animated talking person looks natural. 1 Introduction Research on video and speech processing is usually done independently. However, there is a high correlation between both modalities that is exploited in applications like visual speech recognition and speech driven lip motion synthesis. In the latter the animation of the lips can typically be derived either f...