Stemming versus multi-words indexing for Arabic documents classification

Mohamed Salim El Bazzi, Taher Zaki, Driss Mammass, Abdelatif Ennaji · 2016

Documents indexing is the main step in a conventional document classification or information retrieval framework. This study aims to highlight the influence of features' type on the efficiency of a classification system. Empirical results on Arabic dataset reveal that the choice of extracted feature's type has a significant impact on conserving semantic information and improving classification accuracy, especially with the morphological complexity of the Arabic language. Precision, recall, F-measure and accuracy are the metrics adopted to compare the efficiency of the proposed indexing system.

Read the paper · More papers on PaperTik