SD-CSMOTE: Over-sampling method based on SNN-DPC and improved SMOTE
He Ma, Xu Zhang, Mei Song, Yi Zhu, Wei‐Chiang Hong · Neurocomputing · 2024
An over-sampling method, SD-CSMOTE, is proposed to address the problem of intra-class imbalance in data. First, the minority samples are clustered via the shared nearest neighbour-density peak clustering method . The sample density within each cluster is then determined using the kernel density estimation method. Thereafter total number of samples to be generated in each cluster is calculated. Finally, two samples are randomly selected from the cluster on the basis of the cluster centre, and a new sample is created at the centroid of the three samples. The proposed method is compared with 10 different sampling methods through extensive tests on three classifiers and 10 publicly imbalanced datasets. The results demonstrate that the proposed method achieves optimal F1 value, G-mean, and AUC on most of the datasets, and the Friedman rankings are observed to be optimal across multiple classifiers . It is confirmed that the proposed method performs better than other sampling methods in resolving the intra-class imbalance problem. This method also provides a new way to address issues such as small disjuncts and data imbalance.