Morphosyntactic Parser and Textual Corpora
Aleksei Dobrov, Anastasia Dobrova, Pavel Grokhovskiy, Nikolay Soms · 2017
This article analyzes the problems of parsing texts with linguistic phenomena of controversial nature which may rarely be encountered in NLP projects focusing on Indo-European languages, but are quite frequent in other languages, e.g. in the corpus of Tibetan Indigenous Grammatical Treatises, therefore, parsing texts with such phenomena is necessary for completeness of automatic morphosyntactic annotation of textual corpora. Development of the morphosyntactic analyzer for the Tibetan language started in 2016 and had already proved to be quite useful to deal with specific phenomena of Tibetan, and with previously unsolvable issues of tokenization. The ultimate goal of the project is to create a consistent formal grammatical description (formal grammar) of the Tibetan language, including all grammar levels of the language system from morphosyntax (syntactics of morphemes) to the syntax of composite sentences and supra-phrasal entities. The previously published version of the automatic morphosyntactic annotation was created on the basis of morphologically tagged corpora of Tibetan texts and had high, but not 100 percent coverage (the ratio of the amount of atoms covered by parse trees to the total amount of atoms), precision and recall. This article describes the problems that had to be solved after that, in order to develop the current version of the morphosyntactic parser which allowed to achieve complete and correct automatic annotation of the corpus, and the chosen ways of solving them, which allowed obtaining a complete morphosyntactic annotation of units previously treated as tokens (lexical tokens, words or other atomic parse elements), but required a substantial refactoring (restructuring existing code without changing its external behavior) of the formal grammar. Thus, not only the frequent, but all the constructions turned out to be important in the construction of the formal model.