Text Mining to Analyze Mammogram Screening Results for Breast Cancer Patients in Saudi Arabia
Mohammed Abdul Salam Gollapalli, Maryam Alqusser, Amal Althobaiti, Lubna Alzaid, Roaa Alorefan, Sara Alnajim, Yasmeen Alsaleem · 2023
Breast cancer (BC) affects women of all ages after puberty around the world, but the rate spontaneously rises as the women get older. Saudi Arabia is home to a large number of BC patients. Many women are diagnosed with BC and are frequently admitted to the hospital, resulting in a massive amount of clinical data that can be analyzed and used for medical research. However, proper text mining algorithms must be implemented in order to extract valuable insights and knowledge discovery from these medical notes. In this study, 7 years of BC patients' medical notes data (between 2012-2018) were officially obtained from the hospital and experimented to diagnose patient issues using three machine learning (ML) algorithms. The dataset contained 62,845 clinical text records on "X-ray screening results" recorded by the doctors after seeing the mammogram for patients who had visited the hospital through emergency and regular appointments. Natural language processing (NLP) models were employed to exploit the text data and converted into symbols. Furthermore, the text clustering algorithms K-means, K-medoids, and agglomerative clustering were carefully experimented. The results indicated that K-means outperforms the other two algorithms, with seven identified clusters. K-means clustering also yielded 0.198 work saved over sampling (WSS) and 5.5 database performance (DB) values and was identified as the best unsupervised learning algorithm for medical text clustering.