Text-to-speech voice adaptation from sparse training data
Alexander B. Kain, Michael W. Macon · 1998
Voice adaptation describes the process of converting the output of a text-to-speech synthesizer voice to sound like a different voice after a training process in which only a small amount of the desired target speaker's speech is seen. We employ a locally linear conversion function based on Gaussian mixture models to map bark-scaled line spectral frequencies. We compare performance for three different estimation methods while varying the number of mixture components and the amount of data used for training. An objective evaluation revealed that all three methods yield similar test results. In perceptual tests, listeners judged the converted speech quality as acceptable and fairly successful in adapting to the target speaker. 1. INTRODUCTION Voice conversion systems aim to modify a source speaker's speech so that it is perceived to be spoken by a different target speaker. Integrating voice conversion technologies into a concatenative speech synthesizer allows for the production of add...