Preprocessing of imbalanced breast cancer data using feature selection combined with over-sampling technique for classification
Janjira Jojan, Anongnart Srivihok · 2013
Class imbalance problems have been found in many medical data in recent years. Data are imbalanced when the distributions of classes are highly imbalanced that means the number of instances of one class is very different to the other classes. Feature selection combined with over-sampling technique (FOT) is proposed to preprocess data before classifying our dataset, imbalanced breast cancer. We used feature selection techniques, Consistency Subset Evaluation, at the beginning to remove insignificant attributes of data. The remaining attributes were fed into over-sampling phase to adjust instances in the minority class. After preprocessed the dataset, we classified data using three classification algorithms, decision tree, BayesNet, and OneR. The f-values of classification data using FOT are 0.76, 0.638, and 0.64, respectively. These are greater than the f-values of three above classifications without FOT as, 0.561, 0.518, and 0.512, respectively. The experimental results indicated that FOT achieves better f-values than non-FOT preprocessing and have performed well in improving the performance of classifiers on this dataset, especially, decision tree.