An LFG Chinese Grammar for Machine Use
Ji Chen Fang, Tracy Holloway King · 2007
This paper describes the Chinese grammar developed at PARC, including its three basic components: the tokenizer and tagger, lexicon and syntactic rules. Some of the challenges and issues that we have encountered in the process of development are discussed. In addition, we present our methods of handling these issues. We also illustrate how we evaluate our grammar, providing the evaluation results and some error analyses. 1 Background Introduction This paper describes a Chinese grammar developed at the Palo Alto Research Center (PARC). This grammar is designed for machine use and is implemented in the framework of Lexical Functional Grammar (LFG) (Kaplan and Bresnan, 1982; Dalrymple, 2001; Bresnan, 2001). LFG is characterized by its two parallel levels of syntactic representation: Constituent Structure (c-structure) and Functional Structure (f-structure). C-structure encodes information about phrasal structure and linear word order. F-structure encodes information about ‘the various functional relations between parts of sentences, information like what is the subject and what is the predicate’ (Sells, 1985). Both c-structure information and f-structure information are carried in syntactic rules such as (1). (1) S –> NP: (ˆ SUBJ) = !; VP: ˆ = !. The ˆ refers to the f-structure of the mother node and the ! refers to the f-structure of the node itself. (ˆSUBJ) = ! means that the SUBJ part of the mother’s f-structure (the f-structure of the S in (1)) is the f-structure of the node itself (the f-structure of the NP in (1)). ˆ= ! means that the f-structure of the node itself (the VP in (1)) goes into the f-structure of its mother node (the S in (1)); that is, VP is the functional head of S. (2) shows two additional phrase structure rules. Together with the rules in (1), these will derive the c-structure and f-structure in (4) and (5) for example (3). (2) NP –> N: ˆ = !; VP –> V: ˆ = !. †Fuji Xerox funded the initial research on the Chinese grammar described in this paper, and we are especially grateful to the support we have received from Tomoko Ohkuma and Hiroshi Masuichi of Fuji Xerox throughout the development. Professor Bing Swen and Professor Shiwen Yu of Beijing University have provided substantial support for the tokenizer and tagger used in this grammar. We also appreciate the feedback they provided during our conversations regarding ways to improve the tokenizer and tagger. We would also like to thank Yuqing Guo from DCU for her work in developing the gold analyses for the 200 gold sentences against which we evaluate our grammar. We also owe our thanks to Emily M. Bender for her helpful feedback and comments on this paper.