Phonetisaurus-based letter-to-sound transcription for Standard Arabic

El-Hadi Cherifi, Mhania Guerti · 2017

Every TTS system must be able to convert graphemic strings into phonological representations for the purpose of pronouncing the input text. Phonetisaurus is a WFST (weighted finite-state transducer)-driven grapheme-to-phoneme (g2p) framework suitable for rapid development of high quality g2p systems. In this paper we present this fully data-driven technique, the preliminary investigations of the applicability of this advanced technique for automated transcription of Arabic text, and discuss the results and experiments of this method apply on a pronunciation dictionary of the Most Commonly used Arabic Words (MCAW-Dict.). Experimental results illustrate the speed and accuracy of the proposed system.

Read the paper · More papers on PaperTik