Malayalam TTS Using Tacotron2, WaveGlow and Transfer Learning
Aloyise Biju Mathew, Anandu Sunil Kumar, Febin K Dominic, Glen Vipin Pereira, Juby Mathew · 2023
Text-to-speech (TTS) systems have become an essential component in various applications, such as virtual assistants, audiobook production, and accessibility tools. However, developing TTS systems for languages with limited resources and unique phonetic characteristics poses significant challenges, especially for low-resource languages like Malayalam. This paper proposes a novel Malayalam Text-To-Speech system that uses Tacotron2, an established neural network architecture that maps character embeddings to mel-spectrograms, and WaveGlow, a flow-based network that takes mel-spectrograms as input and generates high-quality speech comparable to that of a natural human voice. Due to the scarcity of labelled data for Malayalam, transfer learning is employed to overcome the limitations of insufficient training data. Fine-tuning is performed using a smaller dataset specifically tailored for Malayalam, ensuring better language-specific representation and improved performance. To evaluate the proposed system, Mean Opinion Score (MOS) was used to assess the quality and similarity of synthesized speech to natural speech.