A Novel Algorithm to Extract Tri-Literal Arabic Roots
Mohanned Momani, Jamil Faraj · 2007
Stemming role and root extraction in the context of information retrieval systems is significant particularly for the Arabic language. In this article, we proposed and implemented a novel algorithm to extract tri-literal Arabic roots. Rootless words are filtered out then prefixes and suffixes removal is performed. Double letters that belong to the Arabic word are removed after sorting term letters. Letter removal is conducted until three letters are remained. Finally, the remaining letters are arranged according to their order in the original word. The implementation of the algorithm has been tested on two types of Arabic text documents. The results of both runs were very promising and satisfactory showing over 73% of accuracy.