Confusion Matrix-Based Feature Selection
Sofia Visa, B. Ramsay, Anca Ralescu, E. VanDerKnaap · 2011
This paper introduces a new technique for feature selection and illustrates it on a real data set. Namely, the proposed ap-proach creates subsets of attributes based on two criteria: (1) individual attributes have high discrimination (classification) power; and (2) the attributes in the subset are complemen-tary- that is, they misclassify different classes. The method uses information from a confusion matrix and evaluates one attribute at a time. Keywords: classification, attribute selec-tion, confusion matrix, k-nearest neighbors; Background In classification problems, good accuracy in classification is the primary concern; however, the identification of the at-tributes (or features) having the largest separation power is also of interest. Even more, for very large data sets (such