Arborest – a VISL-Style Treebank Derived from an Estonian Constraint Grammar Corpus

Eckhard Bick, Heli Uibo, Kaili Müürisep · University of Southern Denmark Research Portal (University of Southern Denmark) · 2004

Treebank creation is a very labor-consuming task, especially if the applications intended include machine learning, gold standard parser evaluation or teaching, since only a manually checked syntactically annotated corpus can provide optimal support for these purposes. There are, however, possibilities to make the annotation process (partly) automatic, saving (manual) annotation time and/or allowing the creation of larger corpora. Whenever possible, existing resources – both corpora and grammars – should be reused. In the case of the Estonian treebank project Arborest, we have therefore opted to make use of existing technology and experiences from the VISL project, where two-stage systems including both Constraint Grammar (CG)and Phrase Structure Grammar (PSG)-parsers have been used to build treebanks for several languages (Bick, 2003 [1]). Moreover, the VISL annotation scheme has been adopted as a standard for tagging the parallel corpus in Nordic Treebank Network. For Estonian, there already exists a shallow syntactically annotated – and proof-read – corpus, allowing us to bypass the first step in treebank construction (CG-parsing). This paper describes how a VISL-style hybrid treebank of Estonian has been semi-automatically derived from this corpus with a special Phrase Structure Grammar, using as terminals not words, but CG function tags. We will analyze the results of the experiment and look more thoroughly at adverbials, non-finite verb constructions and complex noun phrases. The questions we will try to answer are:

Read the paper · More papers on PaperTik