A study on speaker-adaptable multilingual synthesis

Javier Latorre · 2006

This thesis introduces a new method for synthesizing multiple languages with the voice of any speaker, so that for example Japanese speech can be synthesized with the voice of a Russian monolingual speaker. This approach is based on the hypothesis that the average voice created by mixing a sufficient number of speakers is the same for all languages, i.e., the average voice is equivalent to a polyglot speaker. To create such an average voice, we use HMM-based speech synthesis. In our method, first data from multiple speakers of different languages is combined to create a speaker independent (SI) model. In the second step, this model is adapted to a given speaker to create a speaker dependent (SD) model that imitates the voice of that speaker. The adaptation is performed by means of supervised Maximum Likelihood Linear Regression with some minutes of speech data from the target speaker. Using such SD model, any of the languages used to train the SI model can be synthesized with the voice of the target speaker, regardless

Read the paper · More papers on PaperTik