Improving Arabic document categorization: Introducing local stem

Eiman Tamah Al-Shammari · 2010

Stemming is a fundamental step in processing textual data preceding the tasks of text mining, Information Retrieval (IR), and natural language processing (NLP). The common goal of stemming is to standardize words by reducing a word to its base (root or stem), thus can be also considered a feature reduction technique. This paper aims at presenting a new dictionary free, content-based Arabic stemmer and adopts it as a feature reduction (selection) mechanism to study its contribution in improving Arabic text categorization. We employed three stemming mechanisms (root-based, light, and our stemming technique and assessed their performance in text classification exercises for an Arabic corpus to compare and contrast the text mining effectiveness of these Arabic stemming algorithms. The experiments were conducted on a corpus consisting of 2,966 Arabic documents that fall into three categories: cultural, social, and general. The experiment results showed that our stemmer significantly improved text classification accuracy.

Read the paper · More papers on PaperTik