Rapid development of new TTS voices by neural network adaptation
Tijana Delić, Siniša Suzić, Milan Sečujski, Darko Jovan Pekar · 2018
Recent development of parametric speech synthesis based on neural networks (NN) has inspired a range of new techniques for multispeaker speech synthesis. In this paper, a very simple but successful method for creating a new NN-based text-to-speech (TTS) voice with a small amount of data is presented. Speech data from the target speaker is used to adapt a network already trained on source speaker data rather than to train a randomly initialized network. This approach reduces the quantity of target speaker data needed for producing high quality synthetic speech in target speaker's voice. All experiments were carried out on American English databases with both female and male speakers. The results were evaluated both objectively and subjectively and it was shown that adapting parameters of an existing model with just 10 minutes of speech from new speaker produces synthetic speech with quality quite comparable to what can be achieved with a 3-hour database using conventional DNN training. It was also shown that the characteristics of the initial model are not of critical importance for synthetic speech quality.