Information extraction: What have we learned?

Wendy G. Lehnert · Discourse Processes · 1997

Practical natural language processing applications typically confront significant difficulties with lexical ambiguity, knowledge‐engineering bottlenecks, and brittle processing techniques that fail to handle unrestricted input in a robust fashion. However, recent research efforts suggest that these problems can be successfully negotiated for an important class of applications. Information extraction systems operate in restricted discourse domains where they summarize the content of input texts in a highly goal‐oriented manner. Machine learning techniques appear to be a promising foundation for high‐quality information extraction performance, requiring only a training corpus of representative input documents that have been annotated by a domain expert. In this paper we will describe innovative research in information extraction, and speculate a bit about the future of information extraction.

Read the paper · More papers on PaperTik