End-to-End Text-to-Speech Systems in Arabic: A Comparative Study
Mayda Alrige, Omaima Almatrafi, Riad Alharbey, Mashail Alotaibi, Maryah Almarri, Lujain Alghamdi · 2024
This study compares the performance of two popular end-to-end text-to-speech (TTS) systems, the Tacotron and its successor, the Tacotron 2, each used with rival vocoders, the WaveNet and the WaveGlow, respectively. We conducted experiments on Nawar Halabi’s dataset, which contains approximately three hours and forty-two minutes of Arabic speech and qualitatively evaluated the models using the mean opinion score (MOS). The original Tacotron with WaveNet achieved 4.2 on a scale of 5 of a score for mean opinion, thus outperforming Tacotron 2 with WaveGlow in terms of naturalness of speech. We found that crafted text analysis is a crucial step in improving end-to-end TTS for complex languages, such as Arabic. We recommend investing more in text preprocessing as a prior step to accounting for language-specific features such as diacritic, as well as enhancing the voice quality produced through prosody modeling.