Fuzzy Topic Modeling for Medical Corpora

Amir Karami · Digital Repository at the University of Maryland (University of Maryland College Park) · 2015

Unstructured data such as text is an important data type in medical, biomedical, and health domains. The majority of patient records are in non-standard text format which is both a challenge in processing and useful for clinical decision making. One of the challenges for text analysis in the medical domain including the clinical notes and research papers is analyzing large-scale medical documents. As a consequence, finding relevant documents has become more difficult for experts and previous work has also shown unique problems with medical documents for expert such as redundancy issue. Looking for ways to automatically retrieve the enormous amount of medical knowledge has always been an intriguing topic. The massive flow of medical documents including scholarly publications and clinical notes has benefited experts by providing access to a huge amount of text data. However, due to this amount of data, medical experts are finding it increasingly difficult to locate the information of interest. As a consequence, finding relevant documents has become more difficult. Effective text mining systems are needed to extract and exploit not only explicitly stated information but also implied and inferred data. Many powerful methods have been developed in recent years to make the text processing automatic. The first idea was bag-of-words. Using bag-of-words leads to the sparse high dimension problem that has low performance and needs high cost of computation. Dimension reduction techniques, especially topic models, are one of the useful techniques to overcome the problems of bag-of-words. The themes in documents help to retrieve documents on the same topic with and without a query. One of the popular methods to retrieve information based on discovering the themes in the documents is topic modeling. In this research we describe a novel approach in topic modeling using fuzzy clustering. To assess the value of our models, we experiment with the text datasets of medical documents. The quantitative evaluation carried out through document modeling, document classification and document clustering shows that the models produce superior performance to LDA, the most-cited topic model article in Google scholar, indicating that fuzzy set theory can be improved the performance of topic models in medical domain. Our approach handles redundancies and short-length document issues in medical domains, and can discover the number of relations between topics in a corpus. This research contributes to the emerging field of understanding the characteristics of the medical documents and utilizes for them in text analytics.

Read the paper · More papers on PaperTik