Breast Cancer Detection from Imbalanced Clinical Data: A Comparative Study of Sampling Methods
Mahsa Bahrami, Mansour Vali, Hanif Kia · 2023
accurately detecting breast cancer presents a distinctive opportunity for addressing and managing its associated side effects. Collecting patient data is often costly, resulting in imbalanced clinical data, which poses significant challenges for machine learning algorithms. In this paper, we provide a comparative analysis of sampling methods for breast cancer detection. We initially pre-processed clinical data, followed by a comparison of various sampling methods to balance the data. Subsequently, we utilized a support vector machine (SVM) for the classification of malignant and benign breast cancer. Random over-sampling, synthetic minority over-sampling technique (SMOTE), borderline SMOTE, K-means SMOTE, adaptive synthetic sampling, under-sampling majority class, Edited Nearest Neighbor (ENN), Repeated Edited Nearest Neighbor (RENN), near miss, Tomeklink, SMOTEEN, and SMOTETomek were compared and evaluated for Breast Cancer Detection (BCD) from imbalanced clinical data. We also computed feature importance with the eXtreme gradient boosting method that offers an exclusive chance for pathologists in the data processing. We validated a comprehensive examination of BCD through a dataset comprising 569 recordings from the Wisconsin Diagnostic Breast Cancer Data. The best performance was achieved by SMOTEEN, where the accuracy, sensitivity, and specificity were 98.4%, 97.6%, and 98.9%, respectively. It was also found that the mean and worse concave points were more important features for BCD in the WDBC dataset.