A Rule and Template Based Stemming A lgorithm for Arabic Language

Tengku Mohd, Tengku Mohd Tengku Sembok, Belal Mustafa, Abu Ata, Zainab Abu Bakar · International Journal of Mathematical Models and Methods in Applied Sciences · 2011

Stemming is defined as the conflation of all variations of specific words to a single form called the root or stem. Stemming plays a vital role in natural language processing and understanding. As in other languages, there is a need for an effective stemming algorithm for Arabic words. Arabic is a language having a rich and complex morphological word structures and rules. An Arabic stemming algorithm based on morphological rules has been developed, and to enhance its effectiveness, a dictionary of root words is used to determine the right stems. The Arabic stemming algorithm developed by AlOmari is studied and a new algorithm is proposed to enhance the performance. The improvements obtained relate to the order in which the dictionary is lookedup and the order in which the morphological rules are applied. Keywords—Stemming, indexing, information retrieval, natural language processing. I. INTRODUCTION ne of the main modules of a document retrieval system is the text processing and indexing of the input documents to obtain the representation of the documents in the form of indexes. These indexes will be the surrogates to the documents and facilitate the process of retrieving relevant documents with respect to the given query. The process of selecting the representation or index terms constitutes a major operation and technique applied in information retrieval systems. Word stemming is one technique normally applied in the indexing process because it helps in reducing the size of the index terms and also proved to help in improving the degree of relevancy in retrieving documents. The stemming process constitutes word morphological analysis based on the language used in order to get the words' stems to represent the documents as well as to function as indexes to the documents for efficient and effective retrieval.

Read the paper · More papers on PaperTik