Creating Emotional Speech for Conversational Agents
Anh Tuan Do, Scott A. King · 2011
This paper presents an automatic, real-time approach which is capable of creating expressive speech using a set of mathematical models. This approach allows showing emotions in synthetic animated speech in both audio and video. We collect the facial muscle movement data through a tracking system with a high-speed camera, and use that data to create mathematical models for the visual signal. Our emotional model drives muscle parameters to control the shape of the face and prosodic parameters to control the generation of synthetic audio. By applying these models, the expressions can be automatically generated with some optional parameters. We demonstrate the utility of our emotional model by developing a chat system that uses the six universal emotions to create synthetic emotional speech.