Discovering, Ranking and Merging RDF Data Cubes

Sebastian Peter Bayerl, Michael Granitzer · 2017

The RDF Data Cube Vocabulary [1] is a W3C recommendation for publishing multi-dimensional data in a semantic web format. A large and growing number of such data cubes is already available in the linked data cloud. Merging multiple isolated cubes can lead to new insights as known from traditional data warehousing. In order to access and integrate the decentralized cubes, a suitable ranking and discovery mechanism is needed. We propose a cube discovery approach, based on the structure definition of the cubes. Syntactic and semantic properties of the cubes are considered to develop a similarity measure for cubes. We present a graph-based shortest-path algorithm that utilizes the DBpedia category dataset and a Word2Vec [2] model, to carry out the discovery process. The structural mapping for the data cubes generated during our ranking process can be leveraged to support the cube integration process. We use a machine learning approach to combine the implemented similarity measures. The evaluation shows, that this combination produces the best results and that the usage of the Hungarian algorithm [3] is suitable to find good structural mappings.

Read the paper · More papers on PaperTik