SMOTE-LMKNN: A Synthetic Minority Oversampling Technique Based on Local Means-Based k-Nearest Neighbor
Shuang Liu · International Journal of Pattern Recognition and Artificial Intelligence · 2022
Traditional classifiers are trapped by the class-imbalanced problem due to the fact that they are biased toward the majority class. Oversampling methods can improve imbalanced classification by creating synthetic minority class samples. Noise generation has been a great challenge in oversampling methods. Filtering-based and direction-change methods are proposed against noise generation. Yet, the adopted noise filters in filtering-based methods are biased to the majority class. Besides, the [Formula: see text]-nearest neighbor (KNN)-based interpolation in filtering-based and direction-change methods is susceptible to abnormal samples (e.g. outliers, noise or unsafe borderline samples). To overcome noise generation while solving the above shortcomings of filtering-based and direction-change methods, this work presents a new synthetic minority oversampling technique based on local means-based KNN (SMOTE-LMKNN). In SMOTE-LMKNN, the local mean-based KNN (LMKNN) is first introduced to describe the local characteristic of imbalanced data. Second, a new LMKNN-based noise filter is proposed to remove noise and unsafe borderline samples. Third, the interpolation between a base sample and its LMKNN is proposed to create synthetic minority class samples. Empirical results of extensive experiments with 18 data sets show that SMOTE-LMKNN is competent compared with seven popular oversampling methods in training KNN classifier and classification and regression tree (CART).