Adding more languages improves unsupervised multilingual part-of-speech tagging
Benjamin Snyder, Tahira Naseem, Jacob Eisenstein, Regina Barzilay · 2009
We investigate the problem of unsupervised part-of-speech tagging when raw parallel data is available in a large number of languages. Patterns of ambiguity vary greatly across languages and therefore even unannotated multilingual data can serve as a learning signal. We propose a non-parametric Bayesian model that connects related tagging decisions across languages through the use of multilingual latent variables. Our experiments show that performance improves steadily as the number of languages increases.