Ontology Population from the Web: An Inductive Logic Programming-Based Approach

Rinaldo J. Lima, Bernard Espinasse, Hilário Oliveira, Fred Freitas · 2014

The rapid growth of the Web and the information overload problem demand the development of practical information extraction (IE) solutions for web content processing. Ontology Population (OP) concerns both the extraction and classification of instances of the concepts and relations defined by an ontology. Developing IE rules for OP is an intensive and time-consuming process. Thus, an automated mechanism, based on machine-learning techniques, able to convert textual data from web pages into ontology instances may be a crucial path. This paper presents an inductive logic programming-based method that automatic induces symbolic extraction rules, which are used for populating a domain ontology with instances of entity classes. This method uses domain-independent linguistic patterns for retrieving candidate instances from web pages, and a WordNet semantic similarity measure as background knowledge to be used as input by a generic inductive logic programming system. Experiments were conducted concerning both the instance classification problem and a comparison with other popular machine learning algorithms, with encouraging results.

Read the paper · More papers on PaperTik