Considerations on speaking style and speaker variability in speech synthesis

Lennart Nord, Björn Granström · The Journal of the Acoustical Society of America · 1991

In the exploration of speaking style and speaker variability, a multispeaker database and a speech production model is used. The structure of the database, which includes professional as well as untrained speakers, makes it possible to extract relevant information by simple search procedures. In perceptual studies both F0 and duration has had an indisputable effect on prosodics but the role of intensity and of segmental variation has been less dear. This has resulted in an emphasis on the former attributes in current speech synthesis schemes. Intensity has a dynamic aspect, discriminating emphasized and reduced stretches of speech. A more global aspect of intensity must be controlled when an attempt is made to model different speaking styles. Specifically, attempts have been made to model the continuum from soft to loud speech. Systematic variation in speech synthesis has been used as a tool to explore possible speaker dimensions, among them reduced and over-articulated speech. Listening experiments have been carried out with the aim to investigate whether it is possible to describe synthesis samples according to different attitudinal and emotional dimensions.

Read the paper · More papers on PaperTik