CLEANSE – Cluster-based Undersampling Method

Małgorzata Bach, Paulina Trofimiak, Daniel Kostrzewa, Aleksandra Werner · Procedia Computer Science · 2023

Class imbalance is a common problem with datasets relating to various areas of life. It causes many traditional machine learning algorithms to tend to misclassify minority samples as majority ones. Despite various studies, the class imbalance still remains a relevant problem for which no one-size-fits-all solution has been found. In this paper, an undersampling method based on clustering is presented. In the proposed approach K-means algorithm is used to cluster data. In homogenous "majority" clusters, i.e., clusters containing objects of only the majority class, objects within the specified distance from the center are removed. In the case of non-homogeneous clusters, objects located at the class decision boundary are removed using the KNN algorithm. As tests have shown, the clustering-based solution can improve classification quality. The results of experiments show that in many cases the proposed solution outperformed other undersampling techniques described in the literature.

Read the paper · More papers on PaperTik