Clinical Information Extraction Using Word Representations

Shervin Malmasi, Hamed Hassanzadeh, Mark Dras · 2015

A central task in clinical information ex-traction is the classification of sentences to identify key information in publications, such as intervention and outcomes. Sur-face tokens and part-of-speech tags have been the most commonly used feature types for this task. In this paper we eval-uate the use of word representations, in-duced from approximately 100m tokens of unlabelled in-domain data, as a form of semi-supervised learning for this task. We take an approach based on unsuper-vised word clusters, using the Brown clus-tering algorithm, with results showing that this method outperforms the standard fea-tures. We inspect the induced word rep-resentations and the resulting discrimina-tive model features to gain further insights about this approach. 1

Read the paper · More papers on PaperTik