Development of the Truku Language Text-to-Speech Model and Its Application in Digital Audiobook Conversion
Yi-Hao Hsiao, Y.H. Guo, Meng-Chi Huang, Wen‐Yi Chang · 2025
The Truku language is one of the minority Indigenous languages of Taiwan. This paper presents the development of Taiwan’s first-ever Truku language text-to-speech (TTS) model, utilizing over 8 hours of high-quality audio recordings collected from professional male and female Truku speakers. The recordings were processed using VITS2 (Variational Inference for Text-to-Speech) technology to train a robust Truku-TTS model, which can accurately synthesize natural sounding spoken language from written Truku text. To facilitate the model training, we leveraged the computational power of the National Center for High-performance Computing (NCHC), we were able to significantly enhance model training efficiency and performance. Furthermore, we integrated Optical Character Recognition (OCR) technology into the Truku-TTS model workflow, allowing us to convert printed Truku language books into digital audiobooks. This innovative approach not only preserves Truku language but also broadens its accessibility, allowing a wider audience to engage with Truku content. The Truku-TTS model contributes substantially to ongoing efforts to preserve and revitalize Truku languages and provides a crucial tool for cultural and linguistic education. Our methodology and results demostrate a framework that can effectively preserve Truku language materials, enhance Truku language promotion services, and further extend these efforts ro cover all 16 Indigeous tribes comprising 42 dialects in Taiwan, achieving the goal of revitalizing and sustaining Indigenous languages.