A Generative Modelling-based Approach for Text-to-Speech Synthesis in Romanian Language
Marincaş Casian, Czibula Gabriela · 2023
In recent decades, technological advancements in the field of Artificial Intelligence (AI) have opened up new opportunities for generating audio content based on text. In particular, Text-to-Speech (TTS) transformation models have become the most popular method for generating speech from written text. These models can be utilised in applications such as audiobooks, various virtual assistants, and video games, and their performances continue to improve, with generated results becoming increasingly close to human voice. However, there are still many challenges regarding generating as natural voice as possible and appropriate expression of emotions. In this paper we will focus on TTS synthesis, understanding the state-of-the-art and exploring possibilities of training these models for languages other than English, diverse voices, and improving the quality of voice generation. We introduce an approach for Romanian language TTS synthesis using generative deep learning models. Our approach yielded promising results on Romanian language datasets, demonstrating the effectiveness of our methodology.