Simple Semi-Supervised POS Tagging

Karl Stratos, Michael Collins · 2015

We tackle the question: how much supervision is needed to achieve state-of-the-art performance in part-of-speech (POS) tagging, if we leverage lexical representations given by the model of Brown et al. (1992)?It has become a standard practice to use automatically induced "Brown clusters" in place of POS tags.We claim that the underlying sequence model for these clusters is particularly well-suited for capturing POS tags.We empirically demonstrate this claim by drastically reducing supervision in POS tagging with these representations.Using either the bit-string form given by the algorithm of Brown et al. (1992) or the (less well-known) embedding form given by the canonical correlation analysis algorithm of Stratos et al. (2014), we can obtain 93% tagging accuracy with just 400 labeled words and achieve state-of-the-art accuracy (> 97%) with less than 1 percent of the original training data.

Read the paper · More papers on PaperTik