Arabtalk, an implementation for Arabic TTS
Yasser Hifny Abdulhalim, Shady Qurany, Salah Eldeen Hamid, Muhsen Rashwan, Muhammad Atiyya, Ahmed Mahmoud, Galaal Khallaaf · 2011
This paper describes the ARABTALK® Text-To-Speech (TTS) synthesis system, developed at RDI , for Arabic language. ARABTALK® is a state-of-the-art corpus based concatenative TTS system. The system employs Artificial Neural Networks (ANN) statistical prosody based models for duration, energy, and global pitch contour prediction. In addition, it has a real time synthesis by selection algorithm to explore large speech corpus. ARABTALK® has a hidden Markov models (HMMs) based procedure to automatically time-align new voices transcriptions to their acoustic phoneme boundaries. In this framework, a mature phonology framework has been developed and many perfect rule based models were utilized in the process of letter to sound conversion. The system is multi-user and safe-threaded enabled for server based applications. This research aims to advance the process of developing high quality Arabic TTS synthesis, which yields natural and human sounding Arabic voices.