State mapping based method for cross-lingual speaker adaptation in HMM-based speech synthesis

Yi-Jian Wu, Yoshihiko Nankaku, Keiichi Tokuda · 2009

A phone mapping-based method had been introduced for cross-lingual speaker adaptation in HMM-based speech syn-thesis. In this paper, we continue to propose a state mapping based method for cross-lingual speaker adaptation, where the state mapping between voice models in source and target lan-guages is established under minimum Kullback-Leibler diver-gence (KLD) criterion. We introduce two approaches to use the established mapping information for cross-lingual speaker adaptation, including data mapping and transform mapping ap-proaches. From the experimental results, the state mapping based method outperformed the phone mapping based method. In addition, the data mapping approach achieved better speaker similarity, and the transform mapping approach achieved better speech quality after cross-lingual speaker adaptation. Index Terms: Speech synthesis, HMM, speaker adaptation, minimum generation error, linear regression

Read the paper · More papers on PaperTik