Synthesis Ensemble Oversampling and Ensemble Tree-Based Machine Learning for Class Imbalance Problem in Breast Cancer Diagnosis

Slamet Sudaryanto N, Mauridhi Hery Purnomo, Diana Purwitasari, Eko Mulyanto Yuniarno · 2022

The Wisconsin Breast Cancer Database dataset describes the imbalanced class. The imbalanced class will produce accuracy that only favors the majority class but not the minority class. Several ensemble oversampling methods are SMOTE and Random Over Sampling. Meanwhile, the tree-based machine learning ensemble used is Random Forest, Adaptive Boosting, and eXtreme Gradient Boosting. At the level 1 ensemble stage, one of the ensemble models with the best performance will be selected as input for the level 2 ensemble process. The level 2 ensemble is a boosting ensemble, where the results of the best ensemble model chosen at the level 1 ensemble will be used as the base model for boosting the XGBoost algorithm. The results were tested with 10 Fold Cross Validation of 0.981, Accuracy 0.987, Recall 0.980 and Precision 0.982. The performance of our proposed framework outperforms several recent classification studies in the breast cancer domain.

Read the paper · More papers on PaperTik