Generating emotional speech with a concatenative synthesizer

Erhard Rank, Hannes Pirker · 1998

We describe the attempt to synthesize emotional speech with a concatenative speech synthesizer using a parameter space covering not only f0, duration and amplitude, but also voice quality parameters, spectral energy distribution, harmonics-to-noise ratio, and articulatory precision. The application of these extended parameter set offers the possibility to combine the high segmental quality of concatenative synthesis with a wider range of control settings needed for the synthesis of natural affected speech. 1 INTRODUCTION The quality of synthesized speech is usually measured in terms of intelligibility and naturalness. State of the art synthesizers are well intelligible under regular conditions. Therefore, development has concentrated on the improvement of naturalness. An obvious way to follow is the incorporation of correct and functionally adequate prosody. Another maybe less obvious but very interesting aim is the generation of non-neutral affect. For the English language a system f...

Read the paper · More papers on PaperTik