Machine Learning of Extraction Patterns from Unannotated Corpora: Position Statement
Roman Yangarber, Ralph Grishman · 2000
One of the principal challenges of information extraction is the efficient customization of a system to a new extraction task. In many cases, extraction is performed by matching the input text against a set of patterns, so the core problem is one of finding the set of patterns for the new task. These patterns may be stated in terms of individual tokens, sequences of basic syntactic constituents such as noun and verb groups, or in terms of syntactic relationships among constituents; the comments here apply equally to any approach, although the mechanisms required will differ. Some extraction tasks may require substantial intersentential inference in addition to intrasentential pattern recognition; we do not address such tasks here.