A dynamical system model for generating fundamental frequency for speech synthesis
K.N. Ross, Mari Ostendorf · IEEE Transactions on Speech and Audio Processing · 1999
Higher quality speech synthesis is required for widespread use of text to-speech (TTS) technology, and prosody is one component of synthesis technology with the greatest need for improvement. This paper describes a new approach to generation of two important cues to prosodic patterns-fundamental frequency (F/sub 0/) and energy contours-given symbolic prosodic labels and text. Specifically, the approach represents vectors of F/sub 0/ and energy with a dynamical system model, which allows automatic estimation of parameters from labeled speech. Parameters at different time scales in the model are structured to capture segment, syllable, phrase and discourse level effects based on linguistic research. F/sub 0/ generation experiments with the dynamical system model show improved synthetic speech quality over the hybrid target/filter approach.