Provisioning and Usage of Provenance Data in the WebIsALOD Knowledge Graph.
Sven Hertling, Heiko Paulheim · 2018
The WebIsALOD dataset provides a linked data endpoint to the WebIsA database, which harvests millions of subsumption relations from a large scale Web crawl using text patterns. For each of the relations, the dataset also contains rich provenance data, such as the text pattern used, the original sentence in which the pattern was found, and the source on the Web. In this paper, we describe several alternatives and design decisions for providing statement-level provenance information at large scale for the WebIsALOD dataset. Furthermore, we show the practical impact of that provenance information for computing confidence scores approximating the correctness of each subsumption relation.