Boosting the creation of a treebank
Blanca Arias-Badia, Núria Bel, Mercè Lorente Casafont, Montserrat Marimon, Alba Milà-Garcia, Jorge Vivaldi, Lluís Padró, Marina Fomicheva, Imanol Larrea · 2014
We present the results of the experiment of bootstrapping a Treebank for Catalan by using a Dependency Parser trained with Spanish sentences.In order to save time and cost, our approach was to profit from the typological similarities between Catalan and Spanish to create a first Catalan data set quickly by (i) automatically annotating with a delexicalized Spanish parser, (ii) manually correcting the parses, and (iii) using the Catalan corrected sentences to train a Catalan parser.The results showed that the number of parsed sentences required to train a Catalan parser is about 1000, which were achieved in 4 months with 2 annotators.