Stemming as a feature reduction technique for Arabic Text Categorization

Fouzi Harrag, Eyas El-Qawasmah, AbdulMalik S. Al‐Salman · 2011

In this paper, a comparative study is conducted for three text preprocessing techniques in the context of the Arabic text categorization problem using an in-house Arabic dataset. We evaluated and compared three Stemming techniques. They are: Light-Stemming, Root-Based-Stemming and Dictionary-Lookup-Stemming. The purpose is to reduce the feature space into an input space of much lower dimension for two different state-of-the art classifiers: Artificial Neural Networks and support vector machines. The results illustrated that using light stemmer enhances the performance of Arabic Text Categorization. The results also showed that the proposed Artificial Neural Networks model was able to achieve high categorization effectiveness as measured by Macro-Average F1 measure.

Read the paper · More papers on PaperTik