Creating large-scale training and test corpora for extracting structured data from the web

Robert Meusel, Heiko Paulheim · MADOC (University of Mannheim) · 2015

For making the web of linked data grow, information extraction methods are a good alternative for manual dataset curation, since there is an abundance of semi-structured and unstructured information which can be harvested that way.At the same time, existing structured data sets can be used for training and evaluating such information extraction systems.In this paper, we introduce a method for creating training and test corpora from websites annotated with structured data.Using different classes in schema.organd websites annotated with Microdata, we show how training and test data can be curated at large scale and across various domains.Furthermore, we discuss how negative examples can be generated as well as open challenges and future directs for this kind of training data curation.

Read the paper · More papers on PaperTik