Issues in training the TreeTagger for Georgian
Sophiko Daraselia · Corpora · 2024
The paper describes the process of retraining the TreeTagger program ( Schmid, 1994 ) for the Georgian language. This includes considering some general procedures such as designing a training corpus, creating a tagging lexicon, and training the TreeTagger on Georgian texts. I use a novel katag tagset and enclitic tokenisation approach in part-of-speech tagging. The katag tagset is based on a new morphosyntactic language model (Daraselia and Hardie, forthcoming). In this paper, I address some major disambiguation considerations that were revealed when training the TreeTagger on Georgian texts. I discuss some ways to get around these matters, such as implementing a workaround to the tagging lexicon. I report on the performance of the TreeTagger program and compare how different parameters such as the size of the training lexicon or context and affix lengths affect the Tagger’s performance.