Pronominal Resolution in Automatic Text Summarisation

Martin Hassel · 2000

Automatic text summarisation is the technique where a computer program summarises a text. One summarisation technique is to, on linguistical and statistical grounds, extract sentences that are central to the text and use them to form a shorter text, the summary. In this case, automatically summarised text can sometimes result in dangling anaphors (i.e. broken anaphoric references). This is due to the fact that the sentences are extracted without making any deeper linguistic analysis of the text. We will in this thesis show a novel method to resolve some types of pronouns and to summarise the text such that coherence and important information is preserved. The method is implemented as a text pre-processor, called PRM (Pronoun Resolution Module) written in Perl. PRM uses lists of likely focuses (Sidner 1984), here called focus applicants. These lists are by order of insertion sorted in order of likelihood. Choice of applicant for an anaphor is based upon salience (represented by the antecedents position in a list) and semantic likelihood (based on what list the antecedent is to be found in). The latter is determined by using semantic information in a noun lexicon. This method enables us to resolve anaphors non-linearly. The domain is Swedish HTML-tagged newspaper text. A live version of the automatic summariser SweSum and PRM can be found at http://www.nada.kth.se/~xmartin.

Read the paper · More papers on PaperTik