Simple Semi-Supervised Training of Part-Of-Speech Taggers

Anders Søgaard · 2010

Most attempts to train part-of-speech taggers on a mixture of labeled and unlabeled data have failed. In this work stacked learning is used to reduce tagging to a classification task. This simplifies semisupervised training considerably. Our prefered semi-supervised method combines tri-training (Li and Zhou, 2005) and disagreement-based co-training. On the Wall Street Journal, we obtain an error reduction of 4.2 % with SVMTool (Gimenez and Marquez, 2004). 1

Read the paper · More papers on PaperTik