Imbalanced data classification scheme based on G-SMOTE

Shoulei Lu, Jun Ye · Procedia Computer Science · 2024

The imbalanced data problem is essential in data identification and classification tasks. Many solutions solve this problem through different methods, but the most common method is to increase the sample size through various sampling methods. The SMOTE method and its related improved methods are relatively effective, but noisy samples are easily generated during interpolation. The G-SMOTE method solves the noise problem to a certain extent. Still, this method divides the data set into two categories: the majority class and the minority class. This approach doesn't work in the real world. To address this problem, we improve the G-SMOTE method, classify the data into clusters based on different feature classes of the data., identify each minority class data set in the data set according to the imbalanced coefficient IR, and then define the decision boundary and hypersphere, and generate diverse samples through different selection strategies. Our solution solves the noise problem and the multi-category imbalance in the data set.

Read the paper · More papers on PaperTik