Voice Conversion Algorithm Based on Gaussian Mixture Model Applied to STRAIGHT

Tomoki Toda, Jinlin Lu, Satoshi Nakamura, Kiyohiro Shikano · Institutional Repositories DataBase (IRDB) · 2000

Voice conversion is a technique used to convert one speaker's voice into another speaker's voice. As a typical voice conversion algorithm, the codebook mapping method has been studied by Abe et al. The main shortcoming of this method is the fact that the acoustic space of a speaker is limited to a discrete representation. To represent the acoustic space continuously, the algorithm based on the Gaussian mixture model (GMM) has also been proposed by Stylianou et al. In this paper, we apply this GMM-based voice conversion algorithm to STRAIGHT proposed by Kawahara et al., which is recognized as a high quality vocoder. In order to evaluate this voice conversion algorithm, we performed subjective and objective experiments on speaker individuality and speech quality, comparing with the method based on the codebook mapping. As results, the performance of the GMM-based voice conversion algorithm is better than that of the codebook mapping method. Effects by the amount of training data for the voice conversion algorithms were also investigated, as well as the number of the Gaussian mixtures. These evaluation results clarify that the GMM-based voice conversion algorithm is successfully applied to STRAIGHT.

Read the paper · More papers on PaperTik