Building a Turkish Text-to-Speech Engine: Addressing Linguistic and Technical Challenges
Tuğçe Melike Koçak, Mehmet Büyükzincir · 2023
The pursuit of end-to-end text-to-speech (TTS) systems has resulted in generating natural sounding speech. However, due to the requirements of high quality data and pronouncing dictionary, certain languages may face obstacles in achieving comparable results. In this paper, we present innovative solutions to these challenges in the context of Turkish TTS applications. To demonstrate the effectiveness of our approach, we train Fastspeech2 and Tacotron models and evaluated them using mean opinion score (MOS) measure and Speaker Encoder Cosine Similarity (SECS). Evaluation results show that FastSpeech2 model performed well than Tacotron, also our approaches for the dataset is improved the performance of FastSpeech2.