Mining Text for Disease Diagnosis

Shusaku Tsumoto, Tomohiro Kimura, Haruko Iwata, Shoji Hirano · Procedia Computer Science · 2017

Electronic patient records (EPR) are rich in texts, where almost all the decision making processes of medical staff are written. Thus, mining in EPR is important for acquision of decision making process and diagnosis. In this paper, as a first step, we focus on text mining for discharge summaries, which include the compact explanation for the patient’s admission. A record of her complaints, physical findings, laboratory results and radiographic studies while hospitalized; a list of changes in her medications at discharge; and recommendations for follow up care. Text mining process consists of the following four processes: first, morphological analysis is applied to a set of summaries and a term matrix is generated. Second, correspond analysis is applied to the classification labels and the term matrix and generates two dimensional coordinates. By measuring the distances between categories and the assigned points, ranking of key words will be generated. Then, keywords are selected as attributes according to the rank, and training examples for classifiers will be generated. Finally, learning methods are applied to the training examples. Experimental validation shows that random forest achieved the best performance and the second best was the deep learner with a small difference, but decision tree methods with many keywords performed only a little worse than neural network or deep learning methods.

Read the paper · More papers on PaperTik