Data selection in semi-supervised learning for name tagging

Heng Ji, Ralph Grishman · 2006

We present two semi-supervised learning techniques to improve a state-of-the-art multi-lingual name tagger.For English and Chinese, the overall system obtains 1.7% -2.1% improvement in F-measure, representing a 13.5% -17.4% relative reduction in the spurious, missing, and incorrect tags.We also conclude that simply relying upon large corpora is not in itself sufficient: we must pay attention to unlabeled data selection too.We describe effective measures to automatically select documents and sentences.

Read the paper · More papers on PaperTik