Learning Embedded Discourse Mechanisms for Information Extraction
Andrew Kehler · 1998
We address the problem of learning discourse-level merging strategies within the context of a natural language information extraction system. While we report on work currently in progress, results of preliminary experiments employing classification tree learning, maximum entropy modeling, and clustering methods are described. We also discuss motivations for moving away from supervised methods and toward unsupervised or weakly supervised methods. Introduction In recent years, there have been two distinct but related trends in natural language processing (NLP) research. The first we call the automation trend, characterized by the movement from systems in which complicated data and rules are hand-crafted to those in which they are acquired automatically, through either symbolic or statistical learning. This movement is almost certainly a prerequisite to achieving broad deployment of NLP technology; nonetheless, it continues to be a work in progress, as NLP systems are still predominantly...