Crawl me maybe: iterative linked dataset preservation
Besnik Fetahu, Ujwal Gadiraju, Stefan Dietze · 2014
Abstract. The abundance of Linked Data being published, updated, and interlinked calls for strategies to preserve datasets in a scalable way. In this paper, we propose a system that iteratively crawls and captures the evolution of linked datasets based on flexible crawl defi-nitions. The captured deltas of datasets are decomposed into two con-ceptual sets: evolution of (i)metadata and (ii)the actual data covering schema and instance-level statements. The changes are represented as logs which determine three main operations: insertions, updates and deletions. Crawled data is stored in a relational database, for efficiency purposes, while exposing the diffs of a dataset and its live version in RDF format.