GUM: A Guided Undersampling Method to Preprocess Imbalanced Datasets for Classification

Kisuk Sung, W. Eric Brown, Erick Moreno‐Centeno, Yu Cheng Ding · 2022 IEEE 18th International Conference on Automation Science and Engineering (CASE) · 2022

In imbalanced datasets, where the majority class has significantly more instances than the minority class, conventional classification methods exhibit poor minority-class detection performance because they tend to classify most instances as majority instances. To address this problem, this paper presents a general-purpose imbalanced-data preprocessing method that combines two instance-selecting techniques to obtain a clean and balanced set of training instances. The first technique, ensemble outlier filtering, removes outlier instances from both majority and minority classes. The second technique, normalized-cut sampling, samples the majority class aiming to preserve its distribution across the majority region. Our proposed data preprocessing method uses these two techniques and can be combined with any general classification methodology on the sub-sampled data to construct a classification model. Computational results show the proposed method outperforms several widely used imbalanced-data classification methods.

Read the paper · More papers on PaperTik