Class-dependent feature selection algorithm for text categorization
Rogerio C. P. Fragoso, Roberto H.W. Pinheiro, George D. C. Cavalcanti · 2016
A common approach in text categorization is to represent each word as a feature, however, many of these features are irrelevant. So, dimensionality reduction is an important step to diminish the computational effort and to improve accuracy. This paper presents a filter method for feature selection called Category-dependent Maximum f Features per Document (cMFDR). cMFDR is an extension that improves the idea of the MFDR algorithm. In MFDR, the best features are selected exploring documents that overcome a threshold that is calculated for the whole dataset under evaluation. We show that having only one global threshold is not an optimal strategy since it disregards categories that contain few relevant features, impairing the classification precision. So, cMFDR computes one threshold per category to assure that every category contributes with a different number of features. Moreover, the threshold calculation is not biased by documents with large number of features, unlike MFDR. The experimental evaluation showed the effectiveness of cMFDR on four text categorization benchmarks using three feature evaluation functions and Naïve Bayes Multinomial classifier. cMFDR obtains better or similar results than MFDR in 98% of the cases.