Relational Web Search

Michael Cafarella, Michele Banko, Oren Etzioni · 2006

Facts are naturally organized in terms of entities, classes, and their relationships as in an entity-relationship diagram or a semantic network. Search engines have eschewed such structures because, in the past, their creation and processing have not been practical at Web scale. This paper introduces the extraction graph, a textual approximation to an entity-relationship graph, which is automatically extracted from Web pages. The extraction graph is an intermediate representation that is more informative than a mere page-hyperlink graph but far easier to construct than a semantic network. The paper also introduces TextRunner, a search engine that utilizes this representation to answer complex relational queries that are difficult to answer using today’s search engines or Web Information Extraction (IE) systems. The paper compares TextRunner to a state-of-the-art IE system on list searches, and finds that TextRunner is 40 % more precise, with 11 % better recall than the IE system. Our experiments, computed over a 90-million page corpus and a 227-million node extraction graph, show how TextRunner will scale to billions of pages.

Read the paper · More papers on PaperTik