Construction and analysis of Tibetan Khampa dialect corpus for speech synthesis

Yi Feng Zhu, Wenhuan Lu, Yangzom, Mengfei Hu, Kuntharrgyal Khysru, Jianguo We · 2023

High-quality corpus is the basis for supporting speech synthesis research, and the quality and efficiency of speech synthesis are directly affected by a good or bad speech synthesis corpus. The limited scope of Tibetan language usage and the complexity of phonology and grammar make the resources for this language relatively scarce, leading to a lag in the development of related technology applications. To address this situation, we designed an filtering algorithm based on phoneme balanced Tibetan text data, and selected 9179 Tibetan utterances from a large-scale text database as the text corpus to be used in building the corpus; then, we selected a professional male presenter of the Tibetan Khampa dialect to record the text corpus and build a speech synthesis corpus of about 25 hours in the Tibetan Khampa dialect. The statistical analysis and experimental results show that the synthesis model trained with this corpus can obtain natural and comprehensible speech, which verifies the completeness, validity and usefulness of this corpus.

Read the paper · More papers on PaperTik