Evaluating Cervical Cancer Risk Using Machine Learning
Tugba Muhlise Okyay, Ibrahim Yilmaz, Macit Koldas · Medical Bulletin of Haseki · 2025
an essential preventive measure supported by global health authorities (4-6).Early detection is critical to reducing mortality, yet asymptomatic progression in the early stages makes timely diagnosis (Dx) challenging (7).Traditional screening methods such as Pap smears and HPV tests, while effective, may be limited in sensitivity, accessibility, or cost in certain healthcare settings.For example, Pap smears may yield false-negative results in up to 50% of cases, leading to delayed diagnosis (8).In addition, in many lowand middle-income countries, limited access to trained Abs tractAim: Cervical cancer development is influenced by a complex interaction of socio-demographic, behavioral, and clinical factors, which can be systematically analyzed using large datasets.Therefore, this study aimed to evaluate the effectiveness of machine learning (ML) models applied to the University of California, Irvine (UCI), cervical cancer risk factors dataset in predicting cervical health outcomes and supporting early detection strategies.Methods: This study was designed as a retrospective data analysis covering a random sampling of patients between 2012 and 2013 who attended the gynecology service at Hospital Universitario de Caracas in Caracas, Venezuela.The publicly available UCI cervical cancer risk factors dataset was utilized for the analysis.A correlation heatmap was generated to explore the relationships among various risk factors.To address the class imbalance present in the dataset, the synthetic minority over-sampling technique (SMOTE) was applied.Subsequently, different ML classifiers were trained and evaluated to predict cervical cancer outcomes with improved accuracy.Results: The correlation analysis revealed strong correlations among smoking-related measures and diagnostic variables, indicating internal consistency.After applying SMOTE, the dataset achieved a balanced distribution of healthy and diseased individuals.The ensemble classifiers demonstrated high accuracy, up to 97%, and precision, with random forest and light gradient boosting machine performing particularly well.However, the recall for cancer detection was lower: 0.80, indicating potential missed diagnoses. Conclusion:The findings support the integration of ML in clinical diagnostics for cervical cancer, highlighting its potential for improving early detection and patient outcomes while also emphasizing the need for ongoing refinement in model performance.