Bootstrapping a Swedish Treebank Using Cross-Corpus Harmonization and Annotation Projection
Joakim Nivre, Beáta Megyesi · DSpace repository (University of Tartu) · 2007
In this paper, we describe an ongoing project with the aim of boot-strapping a large Swedish treebank, ultimately with a size of about 1.5 million tokens, by reusing two previously existing annotated corpora: an old treebank of about 350,000 tokens and a more recently developed part-of-speech-tagged corpus of about 1,2 million words. A key com-ponent in the bootstrapping methodology is the use of cross-corpus harmonization and annotation projection. 1