Fine-Grained Morpho-Syntactic Analysis for the Under-Resourced Language Chaghatay

Kenneth Steimel, Akbar Amat, Arienne M. Dwyer, Sandra Kübler · 2020

We investigate part of speech (POS) tagging for Chaghatay, a historical language with a considerable amount of morphology but few available resources such as POS annotated corpora.In a situation where we have little training data but a large POS tagset, it is not obvious which method will be best to obtain an accurate POS tagger.We experiment with a conditional random field and a Recurrent Neural Network, augmenting the models with coarse grained POS tag information, and by utilizing additional data, either additional unannotated data used to train a language model or annotated data from a modern relative, Uyghur.Our results show that the combination of an RNN and pretraining with coarse grained POS tags reaches the highest accuracy of 76.17%.

Read the paper · More papers on PaperTik