Improving anaphora resolution by identifying animate entities in texts

Richard Evans, Constantin Orǎsan · 2002

Some references to human beings can be identified in English texts using named entity (Chinchor, 1997) and pronoun recognition but in some genres this still leaves a large number of references to people unidentified. The remaining noun phrases have no overt marking as to their animacy and clues as to the appropriate classification of a NP as animate or inanimate are scarce in the surrounding textual material. In English anaphora resolution, recognition of the animacy of NPs improves the accuracy with which gender agreement restrictions can be enforced between pronouns and candidates. In this work, the animacy of English NPs is identified using a combination of a number of tactics. The main one is the use of WordNet. Three noun hierarchies were identified as being indicative of animate entities. The remaining hierarchies are taken to indicate nouns referring to inanimate entities. Four hierarchies of verb senses were identified as containing verbs expected to require an animate subject. By examining the distribution of textual head nouns and main verbs in WordNet, the animacy of the NP or subject NP is assessed. A number of heuristics are used to reinforce or undermine the system’s confidence as to the animacy of a NP. The new system is incorporated into an existing pronominal anaphora resolution system and tested on a text with a high proportion of gendermarked pronouns. The performance of the system using the new method for animate entity recognition is compared with that of the original. 1. Defining the problem Most approaches to pronominal anaphora resolution rely on compatibility of the agreement features between pronouns and antecedents. Although, as noted in (Barlow, 1998), this assumption does not always apply, it is reliable in enough cases to be of great practical value in anaphora resolution systems. Such systems rely on knowledge about the number and gender of noun phrase (NP) candidates in order to check the compatibility between pronouns and candidates (e.g. Hobbs, 1976; Lappin and Leass, 1994; Nasukawa, 1994; Kennedy and Boguraev, 1996; Mitkov, 1998). None of the algorithms proposed by these researchers include a method for actually identifying the instantiated values of the NP agreement features. Information as to the number feature of a NP’s referent is usually available to systems as a result of the preprocessing phase. Identification of NP gender is not so trivial, though numerous researchers (Hale and Charniak, 1998; Denber, 1998; Cardie and Wagstaff, 1999) have proposed automatic methods for identifying the potential gender of NPs’ referents. An additional paper reporting the use of the system developed in (Hale and Charniak, 1998) was (Ge et al., 1998). In the cases where an algorithm applies preferences to a set of candidates for antecedent of a pronoun, minimizing the size of the set makes a great contribution to the effectiveness of a system. As noted above, when dealing with pronominal anaphora in English, the identification of number information is easy because the majority of pre-processing programs (part of speech taggers and parsers) associate number information with the nouns that are tagged. The reason for this is that enforcing number agreement between constituents increases the accuracy of these tools. By contrast, most English sentence analysis methods gain nothing in performance by incorporating gender information into the algorithm. In fact, it has been observed that incorporation of gender information into the tag set reduces the accuracy of part of speech taggers for English. Given that manual annotation is an expensive procedure, and different researchers consider that a corpus containing gender information brings no additional benefit, most of the corpora and resulting pre-processing software for English do not associate gender features with nominal expressions. This omission is seen to be of little consequence in the syntactic analysis of sentences, but in the field of pronominal anaphora resolution, the enforcement of number agreement alone is not sufficient as a means of producing suitable sets of candidates for pronouns. This means that other sources must be used in order to derive the information required for gender agreement. In (Cobuild, 1995), it is written that “Something that is animate has life...” We therefore use the term animate to describe entities in texts that may also be referred to using gender-marked pronouns. This set of entities includes people and animals. One step toward the identification of NP gender is the identification of reference to animate entities by NPs in a text. There are a number of wellestablished tactics that may be applied in order to do this. The most reliable clue that an animate entity is being referred to is the use of gender-marked pronouns such as he, him, his, himself, she, her, hers, and herself. Another tactic is the performance of named entity recognition, which has been tackled by numerous researchers, including (Wakao et al., 1996; Mikheev et al., 1999) in the Message Understanding Conferences (MUC) (Chinchor, 1997). The aim of named entity recognition is to correctly classify sequences of capitalised words or potentially mixed sequences of capitalised and noncapitalised words as organisations, locations, or persons. Evaluation has shown that the accuracy with which systems can perform named entity recognition is above the level of 90% for precision and recall (Mikheev et al., 1999), but only for restricted domains. However, sole reliance on such clues and systems may still, in some genres leave a large proportion of references to animate entities unidentified. Named entity recognition must be coupled with identification of animate reference using NPs headed by a common noun in order to provide sufficient recall in the recognition of references to animate entities. In English, NPs and their head nouns often have no overt gender marking and clues as to their animacy are scarce in the surrounding textual material. The difficulty of the animate entity recognition task with respect to common NPs has led to the proposal of a new identification system in this paper.

Read the paper · More papers on PaperTik