Voice Conversion of Non-aligned Data using Unit Selection

Daniel Erro, Ferran Diego, Antonio Bonafonte · 2006

Voice conversion (VC) technology allows to transform the voice of the source speaker so that it is perceived as the voice of a target speaker. One of the applications of VC is speech-to-speech translation where the voice has to inform, not only about what is said, but also about who is the speaker. This paper introduces the different methods submitted by UPC to the TC-STAR second evaluation campaign. One method is based on the LPC model and the other on the Harmonic+Noise Model (HNM). Unit selection techniques are employed so that the methods no longer require parallel sentences during the training phase. We have applied these methods both to intra-lingual and cross-lingual voice conversion. Results from the TC-STAR evaluation show that the speaker identity is successfully transformed with all the methods. Further work is required to increase the quality of the voice so that it achieve the quality of current TTS voices.

Read the paper · More papers on PaperTik