Online index extraction from linked open data sources
Fabio Benedetti, Sonia Bergamaschi, Laura Po · IRIS UNIMORE (University of Modena and Reggio Emilia) · 2014
Abstract. The production of machine-readable data in the form of RDF datasets belonging to the Linked Open Data (LOD) Cloud is growing very fast. However, selecting relevant knowledge sources from the Cloud, assessing the quality and extracting synthetical information from a LOD source are all tasks that require a strong human effort. This paper proposes an approach for the automatic extrac-tion of the more representative information from a LOD source and the creation of a set of indexes that enhance the description of the dataset. These indexes col-lect statistical information regarding the size and the complexity of the dataset (e.g. the number of instances), but also depict all the instantiated classes and the properties among them, supplying user with a synthetical view of the LOD source. The technique is fully implemented in LODeX, a tool able to deal with the performance issues of systems that expose SPARQL endpoints and to cope with the heterogeneity on the knowledge representation of RDF data. An eval-uation on LODeX on a large number of endpoints (244) belonging to the LOD Cloud has been performed and the effectiveness of the index extraction process has been presented. 1