Sight and sound: generating facial expressions and spoken intonation from context.
Catherine Pélachaud, Scott Prevost · 1994
This paper presents a model for automatically producing prosodically appropriate speech and corresponding facial expression for agents that respond to simple database queries in a 3D graphical representation of the world. This work addresses two major issues in human-machine interaction. First, proper intonation is necessary for conveying information structure, including important distinctions of contrast and focus. Second, facial expressions and lip movements often provide additional information about discourse structure, turn-taking protocols and speaker attitudes ([7], [8], [14], [15]). The intonation generation model is based on Combinatory Categorial Grammar (CCG -- cf. [20]), a formalismwhich easily integrates the notions of syntactic constituency, prosodic phrasing and information structure. Based on the CCG grammar, a simple discourse model and a domain-independent knowledge base, the system produces spoken responses to database queries with appropriate intonation. Given the timings for phonemes and intonational phenomena in the speech wave, we produce precise specifications for generating the lip movements and facial expressions for a graphical model of a human head. Results from our current implementation demonstrate the system's ability to generate a variety of intonational possibilities and facial animations for a given sentence depending on the discourse context. Previous work in the area of intonation generation includes studies by Terken ([21]), Houghton, Isard and Pearson (cf. [11]), Davis and Hirschberg (cf. [6], [10]), and Zacharski et al. ([23]). Benoit et al. ([1]), Brooke ([2]), Cohen et al. ([4]), Hill et al. ([9]), Lewis et al. ([12]) and Terzopoulos et al. ([22]) have worked on lip synchronization with speech. 2 The Implementation