Speaker and style adaptation using average voice model for style control in HMM-based speech synthesis
Makoto Tachibana, S. Izawa, Takashi Nose, Takao Kobayashi · IEEE International Conference on Acoustics Speech and Signal Processing · 2008
We propose a technique for synthesizing speech with desired style expressivity of an arbitrary target speaker's voice. In an MLLR-based speaker adaptation technique for multiple regression hidden semi-Markov model (MRHSMM), the quality of synthesized speech crucially depends on the initial MRHSMM trained from a certain source speaker's data and it is not always possible to synthesize natural sounding speech with a given target speaker's voice. To overcome this problem, we perform simultaneous adaptation of speaker and style from an average voice model. Experimental results show that the proposed technique provides more natural sounding speech than the conventional one with speaker adaptation only.