A survey of part-of-speech tagging approaches applied to K’iche’

Francis Morton Tyers, Nick Howell · 2021

We study the performance of several popular neural part-of-speech taggers from the Universal Dependencies ecosystem on Mayan languages using a small corpus of 1435 annotated K'iche' sentences consisting of approximately 10,000 tokens, with encouraging results: F 1 scores 93%+ on lemmatisation, partof-speech and morphological feature assignment.The high performance motivates a crosslanguage part-of-speech tagging study, where K'iche'-trained models are evaluated on two other Mayan languages, Kaqchikel and Uspanteko: performance on Kaqchikel is good, 63-85%, and on Uspanteko modest, 60-71%.Supporting experiments lead us to conclude the relative diversity of morphological features as a plausible explanation for the limiting factors in cross-language tagging performance, providing some direction for future sentence annotation and collection work to support these and other Mayan languages.

Read the paper · More papers on PaperTik