SDDSMOTE:Synthetic Minority Oversampling Technique based on Sample Density Distribution for Enhanced Classification on Imbalanced Microarray Data
Qikang Wan, Xiongshi Deng, Min Li, Haotian Yang · 2022
Microarray gene expression data contain an unbalanced distribution of data samples among different classes, which poses a challenge to machine learning-based cancer diagnosis. In addition, microarray data consists of small samples and a huge number of genes, which cause the curse of dimensionality. In order to enhance the performance of learning models on imbalanced microarray data, we propose a novel preprocessing method based on the SMOTE, named SDDSMOTE (Synthetic Minority Oversampling Technique based on Sample Density Distribution). The whole preprocessing includes two steps. First, by using a feature selection technology, irrelevant genes are eliminated and obtaining reduced gene data. Second, SDDSMOTE is used to rebalance the reduced data. We performed comprehensive experiments to compare SDDSMOTE with other state-of-the-art Oversampling algorithms using two Support Vector Machine and Logistic Regression on 8 publicly available microarray expression data sets. The experimental results show that SDDSMOTE outperforms compared algorithms in terms of various evaluation criteria, such as Accuracy, F-score, G-mean, and AUC, which indicates its superiority.