AGO-FT: An adaptive guided oversampling based on fast space division and trustworthy sampling space for imbalanced noisy datasets

Yi Deng, Min Wu, Yan Ma · 2024

The prevalent imbalance can cause big data models to favor the common class, leading to difficulties in identifying the minority value class, deteriorating their overall performance and decision-making. Oversampling methods have become the prevalent strategy nowadays by synthesizing the minority instances to balance the data distribution. Nevertheless, most of the current oversampling mechanisms are based on SMOTE, which synthesizes new instances randomly with arbitrarily chosen k-nearest neighbors. Consequently, the current oversampling methods are readily constrained by the optimization of suitable k-nearest neighbor hyperparameters, and the unrestricted blind random selection of k-nearest neighbors and instance synthesis by the current sampling mechanisms can degrade performance. Specifically, synthesizing instances based on universal and unavoidable noise introduces chaotic random generalization. To fill these gaps, an adaptive guided oversampling (AGO-FT) based on fast space division and plausible sampling space is proposed. Firstly, a fast sample space partitioning strategy based on the complete random forest is proposed to derive sample space information adaptively based on dataset specificity. Secondly, a spatial information-based dataset-specific noise detection method is employed to detect anomalous noise in order to prevent further performance degradation. Then, a parameter-free sampling neighborhood with high confidence is derived from the set of sample spaces based on plausible frequency filtering. Finally, a new probability weight based on the degree of spatial confusion is proposed to guide rational instances synthesis and alleviate the blindness of the SMOTE oversampling mechanism. Extensive experimental results demonstrate that AGO-FT is superior to 8 baseline oversampling algorithms on 13 real-world datasets and four classical classifiers.

Read the paper · More papers on PaperTik