The Significance Of Low Frequent Terms in Text Classification
Mayy M. Al-Tahrawi · International Journal of Intelligent Systems · 2014
The significance of low frequent terms in text classification (TC) was always debatable. These terms were often accused of adding noise to the TC process. Nevertheless, some recent studies have proved that they are very helpful in improving the performance of text classifiers. This paper shows the significance of low frequent terms in enhancing the performance of English TC, in terms of precision, recall, F-measure, and accuracy. Six well-known TC algorithms are tested on the benchmark Reuters Data Set, once keeping low frequent terms and another time discarding them. These algorithms are the support vector machines, logistic regression, k-nearest neighbor, naive bayes, the radial basis function networks, and polynomial networks. All the experiments in this research have shown a superior performance of TC when the low frequent terms are used in classification.