Optimizing annotation efforts to build reliable annotated corpora for training statistical models
Cyril Grouin, Thomas Lavergne, Aurélie Névéol · 2014
Creating high-quality manual annotations on text corpus is time-consuming and often requires the work of experts.In order to explore methods for optimizing annotation efforts, we study three key time burdens of the annotation process: (i) multiple annotations, (ii) consensus annotations, and (iii) careful annotations.Through a series of experiments using a corpus of clinical documents annotated for personally identifiable information written in French, we address each of these aspects and draw conclusions on how to make the most of an annotation effort.