SPARQL2Hive: An approach to processing SPARQL queries on Hive based on meta-models

Mouad Banane, Allae Erraissi, Abdessamad Belangour · 2019

The growth of Web data has presented new challenges in terms of the ability to effectively query RDF data. Traditional relational database systems efficiently adapt and query distributed data. With the development of Hadoop, its implementation of the MapReduce Framework with Hive, a data warehouse, the semantics of data processing and querying has changed. We present in this article, SPARQL2Hive a competitive SPARQL Query Processing System on MapReduce that allows ad hoc SPARQL query processing on large RDF graphs. Instead of a direct mapping, SPARQL2Hive uses the query language of Hive, a data warehouse system that queries systems using HDFS, located above Hadoop MapReduce, as an intermediate layer between SPARQL and MapReduce. This additional level of abstraction makes our approach independent of the current version of Hadoop and thus ensures compatibility with future changes to the Hadoop framework as they will be covered by the underlying Hive layer. Our approach is to use the two meta-models of SPARQL and Hive, and propose a transformation between these two meta-models using the ATL language. We compare SPARQL2Hive with MapReduce-based SPARQL implementations.

Read the paper · More papers on PaperTik