Focusing Annotation for Semantic Role Labeling

Daniel Peterson, Martha Stone Palmer, Shumin Wu · 2014

Annotation of data is a time-consuming process, but necessary for many state-of-the-art solutions to NLP tasks, including semantic role labeling (SRL).In this paper, we show that language models may be used to select sentences that are more useful to annotate.We simulate a situation where only a portion of the available data can be annotated, and compare language model based selection against a more typical baseline of randomly selected data.The data is ordered using an off-the-shelf language modeling toolkit.We show that the least probable sentences provide dramatic improved system performance over the baseline, especially when only a small portion of the data is annotated.In fact, the lion's share of the performance can be attained by annotating only 10-20% of the data.This result holds for training a model based on new annotation, as well as when adding domain-specific annotation to a general corpus for domain adaptation.

Read the paper · More papers on PaperTik