Factorizing information extraction from text corpora
Eduard H. Hovy, Donghui Feng · 2007
Automatic information extraction (IE) from unstructured text is critical for solving the information overload problem. As text formats become ever more varied and IE requirements become more demanding, target information becomes harder to define and extract. Therefore, IE procedures become more complex and generally require an iterative cycle involving multiple factors, such as the simple traditional IE framework is no longer adequate. However, a new paradigm of IE to address these issues has not been formalized yet. In this thesis, we develop a framework to formalize a more general and powerful procedure of IE. Our formalization provides a method to rapidly define and structure new IE tasks. For this procedure, we analyze the role and impact of each factor. This thesis makes two kinds of contributions: first, a new high-level and expressive framework that shows the relationships between activities such as domain knowledge modeling and representation, annotation, system building, evaluation, and feedback adjustment, etc., and second, several new approaches to performing various level IE tasks. We start with a simple IE task, extracting biographical facts from the web. Here a traditional, straightforward, and one-pass problem-solving procedure, consisting of definition-learning-testing, is sufficient. Our system automatically learns surface text patterns from dynamic flat web corpora for answering biographical queries. In addition, sentence fragments can serve as knowledge indicators to guide the handling of queries. We then demonstrate the new IE framework using two more-complex tasks. First, we investigate the problem of extracting data records (individual experiments) from the biomedical research literature. In conformance with the elaborated IE framework, we have developed approaches to labeling individual sentences, grouping fields into individual meaningful objects, and scaling the results up to large corpora. We design a novel solution to segment semantic objects based on semantic analysis, for cases where traditional word-similarity-based text segmentation approaches do not work. Second, the IE procedures need be adapted and extended for IE problems emerging from new media formats. We address the task of extracting the most informative message(s) from discussion threads. Since the relationship between thread messages characterizes how knowledge is spread, the IE framework has to accommodate the analysis of source structure. We describe a novel way to classify message and thread topics that requires zero annotation for a supervised approach, using ontological knowledge induced from a canonical text. We also invent a novel HITS-style algorithm with link generation functions to extract the conversation focus of discussion threads.