Improving Text-to-Speech Systems Through Preprocessing and Postprocessing Applications
Saadin Oyucu, Ferdi Doğan · 2023
Text-to-Speech (TTS) systems are significant technologies that provide speech synthesis that is both natural and human-like by converting text to speech signals. TTS systems need every letter, word, and sentence in the text to be correctly pronounced. Therefore, advanced algorithms and artificial intelligence techniques are used to control speech attributes. However, the use of complex models and multilayer deep neural networks alone is not sufficient to synthesize natural speech. In order to enhance the user experience on TTS systems, additional processing is necessary. In addition to steps such as text cleaning, correction and editing, natural language processing (NLP) applications should be integrated into TTS systems. The preprocessing stage improves the quality of the texts and eliminates grammatical errors, resulting in the production of speech expressions with the correct stress and intonation. Moreover, the post-processing of audio signals makes the sound more fluent and emotionally rich. In addition, editing and controlling the attributes employed in speech synthesis will result in a speech experience that is more realistic and user-friendly. This study investigated the effects of various pre-processing and post-processing techniques on Turkish TTS systems for this reason. In the experiments carried out, it was observed how the correction of text input and post-processing of synthesized sounds affect the performance of the TTS system. Pre- and post-processing has been demonstrated to enhance the accuracy, fluency, and emotional expression of TTS systems in experiments. In order to improve the quality of TTS technologies and the user experience, it is suggested that future research should focus on the integration of pre- and post- processing.