Multi-field information extraction and cross-document fusion

Gideon S. Mann, David Yarowsky · 2005

In this paper, we examine the task of extracting a set of biographic facts about target individuals from a collection of Web pages.We automatically annotate training text with positive and negative examples of fact extractions and train Rote, Naïve Bayes, and Conditional Random Field extraction models for fact extraction from individual Web pages.We then propose and evaluate methods for fusing the extracted information across documents to return a consensus answer.A novel cross-field bootstrapping method leverages data interdependencies to yield improved performance.

Read the paper · More papers on PaperTik