Automatic Domain Model Creation Using Pattern-Based Fact Extraction

Christopher Thomas, Pankaj Mehra, Wenbo Wang, Amit Sheth, Gerhard Weikum, Victor Chan · Journal of Bioresource Management · 2010

This paper describes a minimally guided approach to auto-matic domain model creation. The first step is to carve an area of interest out of the Wikipedia hierarchy based on a simple query or other starting point. The second step is to connect the concepts in this domain hierarchy with named relation-ships. A starting point is provided by Linked Open Data, such as DBPedia. Based on these community-generated facts we train a pattern-based fact-extraction algorithm to augment a domain hierarchy with previously unknown relationship oc-currences. Pattern vectors are learned that represent occur-rences of relationships between concepts. The process de-scribed can be fully automated and the number of relation-ships that can be learned grows as the community adds more information. Unlike approaches that are aimed at finding sin-gle, highly indicative patterns, we use the cumulative score of many pattern occurrences to increase extraction recall. The relationship identification process itself is based on positive-only classification of training facts.

Read the paper · More papers on PaperTik