A K-means Clustering Based Under-Sampling Method for Imbalanced Dataset Classification

Chih-Ming Huang, Chuan-Sheng Hung, Yao-Yuan Hsu, You-Cheng Zheng, Cheng-Han Yu, Chun‐Hung Richard Lin, Shi-Huang Chen · 2024

This paper explores the challenges of applying imbalanced datasets to machine learning models. There is a significant disparity in the quantities of different classes, leading to biases in the learning process. Consequently, many models favor predicting the majority class to enhance accuracy, neglecting potentially crucial minority class data. This bias results in an overly optimistic perception of predictive outcomes, leading to erroneous decision-making. This paper proposes a classification sampling method that solves the issue of original imbalanced data. The process is divided into two main parts. The first part involves data preprocessing, utilizing the K-Means algorithm to cluster majority class data. Additionally, it replaces the traditional Euclidean distance with the Hellinger distance as the similarity metric for data clustering. The second part utilizes the representative data obtained in the previous step as training input, thereby reducing the imbalance between different classifications. Finally, experimental results using the imbalanced Kawasaki Disease (KD) dataset, with an imbalance rate of 64.35%, demonstrate that the proposed Hellinger method improves Precision (PPV) by 20.2% compared to the XGBoost method when Recall is above 90%. This effectively addresses the classification bias of traditional learning methods towards imbalanced data, enhancing predictive outcomes for imbalanced datasets. This approach holds potential applications in medical disease prediction, financial fraud detection, and other related domains in the future.

Read the paper · More papers on PaperTik