Experiments in Constructing a Corpus of Discourse Trees
Daniel Marcu, Estibaliz Amorrortu, Magdalena Romera · 1999
We discuss a tagging schema and a tagging tool for labeling the rhetorical structure of texts. We also propose a statistical method for measuring agreement of hierarchical structure annotations and we discuss its strengths and weaknesses. The statistical measure we use suggests that annotators can achieve good levels of agreement on the task of determining the high-level, rhetorical structure of texts. Our empirical experiments also suggest that building discourse parsers that incrementally derive correct rhetorical structures of unrestricted texts without applying any form of backtracking is unfeasible. 1 Introduction Empirical studies of discourse structure have primarily focused on identifying discourse segment boundaries and their linguistic correlates. Very little attention has been paid so far to the highlevel, rhetorical relations that hold between discourse segments. In some cases, the role of these relations was considered to fall outside the scope of a study (Flammia and Zu...