Prosody Analysis and Modeling for Emotional Speech Synthesis

Danning Jiang, Wei Zhang, Liqin Shen, Lianhong Cai · 2006

Current concatenative text-to-speech systems can synthesize varied emotions, but the subtlety and range of the results are limited because large amounts of emotional speech data are required. The paper studies a more flexible approach based on analyzing and modeling emotional prosody features. Perceptual tests are first performed to investigate whether just manipulating prosody features can attain the communication purposes of emotions. Then, based on the positive results, the same corpus, with sufficient prosody coverage, is shared by different emotions in unit selection. Finally, an adaptation algorithm is proposed to predict the emotional prosody features. It models the prosodic variations by linguistic cues and emotional cues separately, and requires only a small amount of data. Experiments on Mandarin show that the adaptation algorithm can obtain appropriate emotional prosody features, and at least several emotions can be synthesized without the use of a special emotional corpus.

Read the paper · More papers on PaperTik