Semi-Supervised Natural Language Processing Approach for Fine-Grained Classification of Medical Reports

Neil Deshmukh · 2019

Although machine learning has become a powerful tool to augment doctors in clinical analysis, the immense amount of labeled data that is necessary to train supervised learning approaches burdens each development task as time and resource intensive. The vast majority of dense clinical information is stored in written reports, detailing pertinent patient information. The challenge with utilizing natural language data for standard model development is due to the complex and unstructured nature of the modality. In this research, a model pipeline was developed to utilize an unsupervised approach to train a language model, a bidirectional recurrent neural network, to generate document encodings; which then can be used as features passed into a classifier model that requires magnitudes less labeled data than previous approaches to differentiate between fine-grained disease classes accurately. The language model was trained on raw clinical reports from the Partner's Healthcare database$(\mathrm{n}=218,159)$and terminated with a loss of 1.62 and a word prediction accuracy of 62%. The classification models were trained on three labeled datasets, occlusion$(\mathrm{n}=1403)$, stroke$(\mathrm{n}=331)$, and hemorrhage$(\mathrm{n}=4350)$, to identify a variety of different occlusions, stroke, and hemorrhages, directly from the clinical report data; resulting in AUCs of 0.98, 0.95, and 0.99, respectively, for the occlusion, stroke, and hemorrhage datasets. The output encodings are able to be used in conjunction with images or waveform data, to create models that can process a multitude of different modalities. The ability to automatically extract relevant features from textual data allows for faster model development and integration of textual modality, overall, allowing clinical reports to become a more viable input for more encompassing and accurate deep learning models.

Read the paper · More papers on PaperTik