A COMPARATIVE STUDY OF THE EFFICIENCY OF DIFFERENT MEASURES TO CLASSIFY ARABIC TEXT

Mohammed Naji Al-Kabi · 2007

The aim of this study is to find the optimal method that can be used to classify Arabic text among the six methods (inner product, cosine, Jaccard, Dice, Naive Bayesian, and Euclidean). Automatic text classification has been needed in many fields for a long time. Many methods are used to classify text. This study will investigate the use of TF-IDF to obtain document vector. A document vector will be used to compute and compare four different associative coefficients of the vector space model (VSM) based on the inner product, cosine, Jaccard and Dice, in order to find the best for Arabic text classification. We found that the cosine measure outperformed the other three associative coefficients of the VSM. Finally we compare the efficiencies of the cosine measure, Naive Bayesian, and Euclidean Measure to classify Arabic text. Experimental results on the same set of Arabic documents used before show that Naive Bayesian slightly outperforms the other methods. Comparison reported in this paper shows that the Naive Bayesian method surpasses the other five methods.

Read the paper · More papers on PaperTik