Uniqueness, Density, and Keyness: Exploring Class Hierarchies

Anja Jentzsch, Hannes Mühleisen, Felix Naumann · Centrum Wiskunde & Informatica (CWI), the national research institute for mathematics and computer science in the Netherlands · 2015

The Web of Data contains a large number of openly-available datasets covering a wide variety of topics.In order to benefit from this massive amount of open data, e.g., to add value to an organization's internal data, such external datasets must be analyzed and understood already at the basic level of data types, uniqueness, constraints, value patterns, etc.For Linked Datasets and other Web data such meta information is currently quite limited or not available at all.Data profiling techniques are needed to compute respective statistics and meta information.Analyzing datasets along the vocabulary-defined taxonomic hierarchies yields further insights, such as the data distribution at different hierarchy levels, or possible mappings betweens vocabularies or datasets.In particular, key candidates for entities are difficult to find in light of the sparsity of property values on the Web of Data.To this end we introduce the concept of keyness and perform a comprehensive analysis of its expressiveness on multiple datasets.

Read the paper · More papers on PaperTik