End-to-End Text-To-Speech synthesis for under resourced South African languages
Thapelo Nthite, Mohohlo Samuel Tšoeu · 2020 International SAUPEC/RobMech/PRASA Conference · 2020
Text-To-Speech (TTS) systems have been widely adopted around the world for various applications, such as reading to the blind and producing speech for dialog systems. There is however, a lack of TTS systems for South African languages due to limited resources. End-to-end TTS methods using Deep Learning have recently been proposed. These methods eliminate the need for time aligned TTS corpora, making them attractive for under resourced languages. In this paper we present the first reported use of an end-to-end approach for implementing TTS systems for isiXhosa and Sesotho. We train the model using the Lwazi II Sotho corpus as well as the Lwazi III Xhosa corpus. The performance of the system is compared to the Qfrency TTS system using a mean opinion score and a word error rate. The results show that the end-to-end system implemented is able to outperform the Qfrency TTS system by an MOS of 0.68 and 0.83 for intelligibility and naturalness respectively. This is one of the first reported implementations of an end-to-end TTS for any South African language.