Machine Learning of Extraction Patterns from Unannotated Corpora: Position Statement

Roman Yangarber, Ralph Grishman · 2000

One of the principal challenges of information extraction is the efficient customization of a system to a new extraction task. In many cases, extraction is performed by matching the input text against a set of patterns, so the core problem is one of finding the set of patterns for the new task. These patterns may be stated in terms of individual tokens, sequences of basic syntactic constituents such as noun and verb groups, or in terms of syntactic relationships among constituents; the comments here apply equally to any approach, although the mechanisms required will differ. Some extraction tasks may require substantial intersentential inference in addition to intrasentential pattern recognition; we do not address such tasks here.

Read the paper · More papers on PaperTik