Discovering Sublanguages in a Large Clinical Corpus through Unsupervised Machine Learning and Information Gain

T. Elizabeth Workman, Guy Divita, Qing Zeng‐Treitler · 2019

Sublanguages are domain-centered subsets of general or colloquial language. Their identification drives several language analysis tasks, but it is difficult to discern separate sublanguages in large clinical corpora. We applied k-means clustering of semantic properties, and a novel implementation of relative entropy as an information gain indicator, to identify sublanguages within a large clinical corpus (~1.6 million documents), visualizing the results in a heat map. Patterns both within and across clusters reveal sublanguage trends. These findings are significant in sublanguage analysis, and have implications on both regional and international levels.

Read the paper · More papers on PaperTik