Realizing Mandarin-Tibetan bilingual speech synthesis by speaker adaptive training

Pei Yu Dong · Journal of Tsinghua University(Science and Technology) · 2013

This paper presents a method to realize hidden Markov model(HMM)-based Mandarin-Tibetan bilingual speech synthesis using the similarities between Mandarin and Tibetan pronunciation.The initial and the final are used as the synthesis units with training using a set of average mixed-lingual models from a large Mandarin multi-speaker-based corpus and a small Tibetan one-speaker-based corpus using speaker adaptive training(SAT).Then,the speaker adaptation transformation is applied to the speaker dependent(SD) training data to obtain a set of speaker dependent Mandarin or Tibetan models from the average mixed-lingual models.The Mandarin speech or Tibetan speech is then synthesized from the speaker dependent Mandarin or Tibetan models.Tests show that this method outperforms the method using only Tibetan SD models when only a small number of Tibetan training utterances are available.

Read the paper · More papers on PaperTik