Unsupervised adaptation of supervised part-of-speech taggers for closely related languages

Yves Scherrer · 2014

When developing NLP tools for low-resource languages, one is often confronted with the lack of annotated data.We propose to circumvent this bottleneck by training a supervised HMM tagger on a closely related language for which annotated data are available, and translating the words in the tagger parameter files into the low-resource language.The translation dictionaries are created with unsupervised lexicon induction techniques that rely only on raw textual data.We obtain a tagging accuracy of up to 89.08% using a Spanish tagger adapted to Catalan, which is 30.66%above the performance of an unadapted Spanish tagger, and 8.88% below the performance of a supervised tagger trained on annotated Catalan data.Furthermore, we evaluate our model on several Romance, Germanic and Slavic languages and obtain tagging accuracies of up to 92%.

Read the paper · More papers on PaperTik