Extracting and Querying a Comprehensive Web Database.
Michael Cafarella · 2009
Recent research in domain-independent information extrac-tion holds the promise of an automatically-constructed struc-tured database derived from the Web. A query system based on this database would offer the same breadth as a Web search engine, but with much more sophisticated query tools than are common today. Unfortunately, these domain-independent Web extractors are usually not model-independent; e.g., an extractor that only finds binary re-lations from text will be blind to relational data found in tables. Because a topic area often has a data model that is a natural fit (e.g., population statistics are usually in ta-bles, while biographical facts about Einstein are embedded in text), even a high-quality domain-independent extractor will miss a substantial amount of data. Our omnivore system attempts to build a comprehen-sive Web database by running multiple domain-independent extractors in parallel over a Web crawl, then combining their outputs into a single large entity-relationship database. Each item in the database describes a single real-world en-tity, and can contain information drawn from a number of popular Web data models. The user can correct flaws in the database, and can query it using either a structured query language or a search-like interface. Due to the Web’s sheer size, users cannot be expected to know the result set’s meta-data a priori, so omnivore automatically chooses an output model and schema when it renders results. In this paper we outline the omnivore architecture and provide specific de-tails about our current prototype.