Synthesizing dialogue speech of Japanese based on the quantitative analysis of prosodic features
Kenzo Hirose, Mayumi Sakata, Hiromichi Kawanami · 2002
Through analyses of the fundamental frequency contours and speech rates of dialogue speech and also of read speech, prosodic rules were derived for the synthesis of spoken dialogue. The fundamental frequency contours were first decomposed into phrase and accent components based on the superpositional model, and then their command magnitudes/amplitudes were analyzed by the method of multiple regression analysis. For the speech rate, the reduction rate of mora duration from reading-style to dialogue-style was calculated. After normalizing the sentence length, the mean reduction rate was calculated as an average over utterances without a complicated syntactic structure. The results of the above analyses were incorporated into prosodic rules for dialogue speech synthesis. Using a previously-developed formant speech synthesiser, synthesis was conducted using both the former rules of read speech and the newly developed rules. A hearing test showed that the new rules can produce better prosody as dialogue speech.