Two dimensional generalization in information extraction

Joyce Yue Chai, Alan W. Biermann, Curry I. Guinn · 1999

In a user-trained information extraction system, the cost of creating the rules for information extraction can be greatly reduced by maximizing the effectiveness of user inputs. If the user specifies one example of a desired extraction, our system automatically tries a variety of generalizations of this rule including generalizations of the terms and permutations of the ordering of significant words. Where modifications of the rules are successful, those rules are incorporated into the extraction set. The theory of such generalizations and a measure of their usefulness is described. Introduction Information extraction (IE) has become a promising area since the advent of the DARPA Message Understanding Conferences (Cowie & Lehnert 1996). Given the vast amount of information available today, successful extraction of useful information has become increasingly important. Most IE systems (MUC6 1995) have used hand-crafted semantic resources for each application domain. However, generation...

Read the paper · More papers on PaperTik