Porting Grammars between Typologically Similar Languages : Japanese to Korean
Roger H. Kim, Mary Dalrymple, Ronald M. Kaplan, Tracy Holloway King · Institutional Repositories DataBase (IRDB) · 2003
The Parallel Grammar project (ParGram) is an international collaboration aimed at producing broad-coverage computational grammars for a variety of languages (Butt et al., 1999; Butt et al., 2002). The grammars (currently of English, French, German, Japanese, Norwegian, and Urdu) are written in the framework of Lexical Functional Grammar (LFG) (Kaplan and Bresnan, 1982; Dalrymple, 2001), and they are constructed using a common engineering and high-speed processing platform for LFG grammars, the XLE (Maxwell and Kaplan, 1993). These grammars, as do all LFG grammars, assign two levels of syntactic representation to the sentences of a language: a superficial phrase structure tree (called a constituent structure or c-structure) and an underlying matrix of features and values (the functional structure or fstructure). The c-structure records the order of words in a sentence and their hierarchical grouping into phrases. The f-structure encodes the grammatical functions, syntactic features, and predicate-argument relations conveyed by the sentence. F-structures are meant to encode a language universal level of analysis, allowing for cross-linguistic parallelism at this level of abstraction. The ParGram project attempts to test the LFG formalism for its universality and coverage and to see how far parallelism can be maintained across languages. Previous ParGram work and much theoretical analysis has largely confirmed the universality claims of LFG theory. The f-structures produced by the grammars for similar constructions in each language have the same major functions and features, with minor variations across languages (e.g., the f-structures for French nouns have a grammatical gender feature but that distinction is not marked in English f-structures). This uniformity has the computational advantage that the grammars can be used in similar applications and that machine translation (Frank, 1999) can be simplified. We have found that it takes roughly two person-years of effort to construct for a new language a grammar that approximates existing grammars in terms of coverage and accuracy (see (Riezler et al., 2002) for a discussion of the coverage and accuracy of the current English grammar). This suggests that the deep-grammar construction task is not as difficult as many people have believed, and indeed may require less effort than would be needed to produce training materials for automatic learning procedures for shallower grammars. However, it is still interesting to explore methods for reducing the linguistic effort that grammar construction requires. To that end, we report here on a preliminary investigation of the difficulty