Creating a test corpus of clinical notes manually tagged for part-of-speech information

Serguei Pakhomov, Anni R. Coden, Christopher G. Chute · 2004

This paper presents a project whose main goal is to construct a corpus of clinical text manually annotated for part-of-speech information. We describe and discuss the process of training three domain experts to perform linguistic annotation. We list some of the challenges as well as encouraging results pertaining to inter-rater agreement and consistency of annotation. We also present preliminary experimental results indicating the necessity for adapting state-of-the-art POS taggers to the sublanguage domain of medical text.

Read the paper · More papers on PaperTik