Fuzzy c-means for extraction and quantification of words representing degrees of diseases (Preprint)
Feng Han, Ziheng Zhang, Hongjian Zhang, Jun Nakaya, Kohsuke Kudo, Katsuhiko Ogasawara · 2022
BACKGROUND Modern medicine generates unstructured data containing a large amount of information. Extracting useful knowledge from this data and making scientific decisions for diagnosing and treating diseases have become increasingly necessary. Unstructured data, such as in the Medical Information Mart for Intensive Care III (MIMIC-III) dataset, contain several ambiguous words demonstrating the subjectivity of doctors. These data can be used to further improve the accuracy of medical support system assessments. OBJECTIVE We propose using fuzzy c-means (FCM) method and Gauss membership to quantify the subjective words in the clinical medical dataset MIMIC-III. METHODS Using 381,091 radiology reports collected from MIMIC-III, we extracted words representing the subjective degree from the text and converted them into corresponding membership intervals based on the words. RESULTS Consequently, the words representing each degree of each disease had a range of corresponding values. Examples of membership medians were atelectasis (2.971), pneumonia (3.121), pneumothorax (2.899), pulmonary edema (3.051), and pulmonary embolus (2.435). These membership sections can determine the symptoms of each disease. CONCLUSIONS In this study, we used the FCM and Gaussian functions to extract words from the MIMIC-III, which represent a subjective degree and cannot be processed by a computer, and performed fuzzy processing on them. It was concluded that words representing the degree in an English interpreted report can be extracted and quantified. The use of these words in medical support systems may improve diagnostic accuracy.