Learning Domain-Specific Discourse Rules for Information Extraction

Stephen Soderland, Wendy G. Lehnert · 1995

This paper describes a system that learns discourse rules for domain-speci#c analysis of unrestricted text. The goal of discourse analysis in this context is to transform locally identi#ed references to relevant information in the text into a coherent representation of the entire text. This involves a complex series of decisions about merging coreferential objects, #ltering out irrelevant information, inferring missing information, and identifying logical relations between domain objects. The Wrap-Up discourse analyzer induces a set of classi#ers from a training corpus to handle these discourse decisions. Wrap-Up is fully trainable, and not only determines what classi#ers are needed based on domain output speci#cations, but automatically selects the features needed by each classi#er. Wrap-Up's classi#ers blend linguistic knowledge with real world domain knowledge. Introduction Discourse analysis takes on a special role in a system that analyzes real-world text such as...

Read the paper · More papers on PaperTik