SMOTE Classification and Random Oversampling Naive Bayes in Imbalanced Data : (Case Study of Early Detection of Cervical Cancer in Indonesia)
Nur Silviyah Rahmi, Ni Wayan Surya Wardhani, Maria Bernadetha Mitakda, Regina Syahla Fauztina, Imelda Salsabila · 2022
Imbalanced data was a problem that is often encountered when classifying, where the distribution of the majority class has more numbers than the minority class. The existence of imbalanced data makes the performance of classification methods in machine learning decrease. This study adopted SMOTE and Random Oversampling (ROS) sampling techniques to overcome imbalanced data combined with the Naive Bayes classification method in cases of detection of early cervical cancer in Indonesia. Cervical cancer is a disease of an abnormal cell group that growth in the cervix (mouth of the womb). Cervical cancer is the most common type and ranks number 2 as cancer suffered by Indonesian women. Various factors that influence the event include eating behavior, personal hygiene behavior, motivational strength, social support, empowerment of knowledge, abilities, and desires. The data used is secondary data with a sample of 72 patients and 20 attributes. A total of21 patients in the classification had cervical cancer and 51 patients did not have cervical cancer. The ratio of 30:70 are imbalanced data. Through this classification method, it is expected to know what factors influence the event of cervical cancer and gains the best performance of two classifications. The results point out that average performance of SMOTE Naive Bayes has a higher (81,73%) than Random Oversampling Naive Bayes which is 81,12%. Therefore, SMOTE Naive Bayes outperforms Random Oversampling Naïve Bayes.