A feature score for classifying class-imbalanced data

Part Pramokchon, Punpiti Piamsa-nga · 2014

Feature ranking method is one of filter-based feature selection which is widely used in text classification. However, many feature scores for ranking produce low classification performance when they are applied to data, where data sizes in each class are drastically different. We present a feature score based on statistical t-test technique, which is a statistical evaluation of the difference between two sample means, to assess the discriminating power of each individual feature. The t-test based feature score can be used to determine whether the numbers of data in each class are drastically unequal. Therefore, the score is insensitive to the problem of class-imbalanced distribution. The multi-class text classification performance of the proposed feature score is compared with seven modern feature scores, which are CMFS, IG, CHI, DF, GINI, OCFS, and DIA. The results show that micro average F1 performance on the Reuters-21578 benchmark dataset by the proposed feature is 94.2%, where of all other metrics are not over 80%.

Read the paper · More papers on PaperTik