Navigating Extracted Data with Schema Discovery.

Michael Cafarella, Dan Mircea Suciu, Oren Etzioni · 2007

Open Information Extraction (OIE) is a recently-introduced type of information extraction that extracts small individual pieces of data from input text without any domainspecific guidance such as special training data or extraction rules. For example, an OIE system might discover the triple Frenzy, year, 1972 from a set of documents about movies. Because OIE is domain-independent, it promises to help users when they have a corpus of structured data, but that structure is unknown, such as when browsing a novel domain or formulating a query. We can describe the structure to the user by displaying a relational schema that fits the extracted data. Unfortunately, the extractions do not carry full schema information: we have extracted values, but not the correct relations, their rows, or their columns. In response we propose TGen, an algorithm for schema discovery, which automatically derives a high-quality relational schema for the extracted data. Different applications have different schema-design requirements, which can be encoded as input to TGen. We show that our data-mining approach runs in minutes on millions of documents while still resulting in schemas that are useful for exploring unfamiliar data or for composing queries over extracted data. 1.

Read the paper · More papers on PaperTik