Parsing with PCFGs and automatic f-structure annotation

Aoife Cahill, Mairéad McCarthy, Josef van Genabith, Andy Way · Dublin City University Open Access Institutional Repository (Dublin City University) · 2002

The development of large coverage, rich unification- (constraint-) based grammar resources is very time consuming, expensive and requires lots of linguistic expertise. In this paper we report initial results on a new methodology that attempts to partially automate the development of substantial parts of large coverage, rich unification-(constraint-) based grammar resources. The method is based on a treebank resource (in our case Penn-II) and an automatic f-structure annotation algorithm that annotates treebank trees with proto-fstructure information. Based on these, we present two parsing architectures: in our pipeline architecture we firstextract a PCFG from the treebank following the method of [Charniak, 1993; Charniak, 1996] , use the PCFG to parse new text, automatically annotate the resulting trees with our f-structure annotation algorithm and generate proto-f-structures. By contrast, in the integrated architecture we firstautomatically annotate the treebank trees with fstructure information and then extract an annotated PCFG (A-PCFG) from the treebank. We then use the A-PCFG to parse new text to generate proto-fstructures. Currently

Read the paper · More papers on PaperTik