Exploring the effects of diacritization on Arabic frequency counts
Osama Hamed, Torsten Zesch · 2018
Statistical natural language processing relies on corpora as a source of token frequency counts. Obtaining reliable frequency counts is challenging in Arabic due to omitted diacritics in almost all writings. In this paper, we explore how severely this situation affects the resulting language models. For this purpose, we analyze the properties of the few available manually diacritized corpora and use them to evaluate automatic diacritization tools. We then apply the best performing tool on non-diacritized texts and investigate the properties of the resulting language models.