Comparative Evaluation of Arabic Language Morphological Analysers and Stemmers
Majdi Sawalha, Eric Atwell · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 2008
Arabic morphological analysers and stemming algorithms have become a popular area of research. Many computational linguists have designed and developed algorithms to solve the problem of morphology and stemming. Each researcher proposed his own gold standard, testing methodology and accuracy measurements to test and compute the accuracy of his algorithm. Therefore, we cannot make comparisons between these algorithms. In this paper we have accomplished two tasks. First, we proposed four different fair and precise accuracy measurements and two 1000-word gold standards taken from the Holy Qur’an and from the Corpus of Contemporary Arabic. Second, we combined the results from the morphological analysers and stemming algorithms by voting after running them on the sample documents. The evaluation of the algorithms shows that Arabic morphology is still a challenge. 1 Three Stemming Algorithms We selected three stemming algorithms for which we had ready access to the implementation and/or results. Shereen Khoja Stemmer: We obtained a Java version of Shereen Khoja’s stemmer (Khoja,1999). Khoja’s stemmer removes the longest suffix and the longest prefix. It then matches the remaining word with verbal and