Gesture-speech interaction in the SmartKom project
Antje Schweitzer · The Journal of the Acoustical Society of America · 2001
The goal of the SmartKom project is to create an intuitive multi-modal dialog system. Information is presented on the display by an artificial lifelike character, combining the output modalities speech, mimics, and gesture. Multi-modality poses new requirements for speech synthesis. Speech may be accompanied by gestures, adding three new aspects: First, synchronization of lip movements; second, temporal alignment of speech and gesture; and third, effects of gestures on prosody. This paper focuses on the last two points. Building on de Ruiter’s (2000) Sketch Model, speech is synthesized independently from the temporal structure of the accompanying gesture. Preparation phase and retraction phase of gestures can be adjusted to align the relevant speech material with the stroke phase. Concerning prosody, results of a pilot study on intonation of deictic elements show that speech material accompanied by gestures is more likely to be prosodically prominent. Gestures in the SmartKom system are always triggered by linguistic properties of the accompanying speech. Most of the gestures are pointing gestures, which are triggered by deictic elements. Such linguistic information is integrated into the concept input to the synthesis module and is taken into account during prosody generation.