Improving summarization through rhetorical parsing tuning

Daniel Marcu · 1998

We study the relationship between the structure of discourse and a set of summarization heuristics that are employed by current systems. A tight coupling of the two enables us to learn genre-specific combinations of heuristics that can be used for disambiguation during discourse parsing. The same coupling enables us to construct discourse structures that yield summaries that contain textual units that are not only important according to a variety of position-, title-, and lexical-similarity-based heuristics, but also central to the main claims of texts. A careful analysis of our results enables us to shed some new light on issues related to summary evaluation and learning. 1 Motivation Current approaches to automatic summarization employ techniques that assume that textual salience correlates with a wide range of linguistic phenomena. Some of these approaches assume that important textual units contain words that are used frequently (Luhn, 1958; Edmundson, 1968) or words that are use...

Read the paper · More papers on PaperTik