Arabic phonetics and phonology for text analytics and natural language processing applications
C Brierley, Majdi Sawalha, Eric Atwell · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 2011
We plan to apply Text Analytics techniques honed on English for corpus-based exploration of Arabic, for Arabic Text-to-Speech (TTS) and other applications. Such techniques depend on a corpus or sample of naturally-occurring language texts capturing empirical data on the phenomena being studied, for example prosodic-syntactic patterns at phrase juncture or perceived pauses in the speech stream. For Modern Standard Arabic, this would require a representative, multi-speaker corpus of transcribed speech with phrase break annotations and ‘gold standard ’ part-of-speech (POS) categories. We can then mine these annotations as well as plain text.