Learning to extract symbolic knowledge from the World Wide Web

Mark W. Craven, Dan DiPasquo, Dayne Freitag, Andrew McCallum, T. M. Mitchell, Kamal Nigam, Seán Slattery · 1998

The goal of the Web-KB project is to develop automatic methods for constructing and maintaining large knowledge bases whose contents mirror those of the World Wide Web. We argue for the feasibility of a system which, given a manually constructed ontology and a seed knowledge base comprising a set of labeled Web pages, learns to instantiate knowledge-base objects and relations from the Web. Such a system could construct a knowledge base supporting concept-oriented queries to the Web, or serve as a resource for Web-based problem solvers and search agents. At the heart of this task are two general sub-problems: identifying instances of the classes, and identifying instances of the relations that are defined in the ontology. The first sub-problem includes the task of recognizing cases in which individual Web pages correspond to the classes of interest, and the second sub-problem includes the task of identifying cases in which pairs of pages instantiate the ontology's relations. We present the results of initial experiments into the use of text classification and relational learning for these tasks, and sketch problems for future research.

Read the paper · More papers on PaperTik