Personalized Taiwanese Speech Synthesis using Cascaded ASR and TTS Framework

Yuan‐Fu Liao, Wen‐Han Hsu, Chen-Ming Pan, Wern-Jun Wang, Matúš Pleva, Daniel Hládek · 2022

To bring endangered Taiwanese language back to life, this paper leveraged a large-scale Taiwanese across Taiwan (TAT) corpus to construct cascaded automatic speech recognition (ASR) and text-to-speech (TTS)-based personalized Taiwanese speech synthesizers to help young people to learn how to speak Taiwanese. This paradigm not only alleviates the low resource, nonparallel corpus and cross-lingual training data problems but also dramatically reduces the fine-tuning data size and training time. Experimental results on a Taiwanese-to-Taiwanese and Mandarin-to-Taiwanese voice conversion tasks had shown that it allows us to successfully produce good personalized Taiwanese TTS with only approximately 3 minutes of data in both cases.

Read the paper · More papers on PaperTik