Issues in training the TreeTagger for Georgian

Sophiko Daraselia · Corpora · 2024

The paper describes the process of retraining the TreeTagger program ( Schmid, 1994 ) for the Georgian language. This includes considering some general procedures such as designing a training corpus, creating a tagging lexicon, and training the TreeTagger on Georgian texts. I use a novel katag tagset and enclitic tokenisation approach in part-of-speech tagging. The katag tagset is based on a new morphosyntactic language model (Daraselia and Hardie, forthcoming). In this paper, I address some major disambiguation considerations that were revealed when training the TreeTagger on Georgian texts. I discuss some ways to get around these matters, such as implementing a workaround to the tagging lexicon. I report on the performance of the TreeTagger program and compare how different parameters such as the size of the training lexicon or context and affix lengths affect the Tagger’s performance.

Read the paper · More papers on PaperTik