Neural Network-based Vocoders in Arabic Speech Synthesis

Zaha Mustafa Badi, Lamia Fathi Abusedra · 2021

Text-to-speech (TTS) system has been used in many applications to generate natural speech that is intelligible for a given text. Thereby, TTS systems can be a very useful tool in many applications. A text-to-speech synthesis consists of a text analysis model, an acoustic model and a vocoder. The modern neural network-based model has significantly improved the speech synthesis quality. However, the English language is often the focus when developing state-of-the-art technologies. Even though Arabic is the formal language for 25 countries, there is no sufficient research or evaluation conducted to study the performance of state-of-the-art TTS approaches on Arabic speech synthesis, and only conventional TTS techniques have been adapted to the Arabic language. In this paper, the two neural-network vocoders Parallel WaveGan and Multi-Band MelGan are implemented. Investigation of the effectiveness of both vocoders for Arabic speech synthesis is carried out. This work involves reconstruction and examination of the selected models; and evaluating the performance and quality of the generated speech. It is found that the Parallel WaveGan outperformed Multi-Band MelGan in terms of Perceptual Evaluation of Speech Quality (PESQ), where the generated speech using Parallel WaveGan scored PESQ of 2.63, while Multi-Band MelGan scored 2.37.

Read the paper · More papers on PaperTik