Feature Selection Based on Two-stage Resampling Technique for Imbalanced Dataset

Dan Zhao, Zhenyi Shen, Shuangxue Zhao · Procedia Computer Science · 2023

The conventional feature selection methods fail to work on imbalanced datasets due to the overfitting issue caused by the extremely imbalanced positive and negative samples. In addition, the large number of negative samples makes the feature selection operation inefficient. To address these issues, an automatic feature selection method for addressing feature selection issues on extremely imbalanced datasets is proposed. An undersampling operation is performed to generate a small balanced dataset at first. Then, the feature importance scores are estimated based on the occurrence of the feature columns in the feature combinations that have a higher score than the base score. Finally, a repeated feature removal operation and classification model retrain are performed on the new dataset generated with the oversampling technique until the removal of the unimportant feature column does not bring any score improvement. In the proposed approach, the oversampling operation can fully utilize the dataset to train the classification model properly and the importance score is easily interpreted as the frequency of each feature column in the satisfied feature combinations. The experiment shows the superior performance of the proposed feature selection method.

Read the paper · More papers on PaperTik