Adaptive boundary oversampling through feature weighting for class imbalance learning
Qiangkui Leng, Zhiyong Li · 2024
Classification is a crucial learning task in machine learning, which fundamentally involves predicting the category of test examples using a classifier generated from a training example set. However, many real-world applications have training sets with imbalanced class distributions, which often hinder the classification performance of learning algorithms. To address this, this paper proposes a feature-weighted adaptive boundary oversampling method. Initially, the method acquires neighbor information of sample points based on weighted Euclidean distance (WED), identifies minority class samples on the boundary according to the distribution of minority class neighbors, calculates the synthetic factor corresponding to boundary samples, and updates the number of samples to be generated for the sample based on its value. Finally, it randomly selects minority class samples from the neighbors to generate new samples according to the synthetic factor. The proposed method is compared with seven sampling methods on a decision tree classifier and 12 imbalanced datasets from KEEL. The results show that the proposed method achieves the best values in F1, G-mean, and AUC (Area under Curve) on most datasets, and has the best Friedman ranking, proving that compared with other sampling methods, it has better performance in handling classification problems in imbalanced data. By setting certain constraints and allocation strategies for the synthetic factor, it can provide ideas for similar research.