RDF Data Storage Techniques for Efficient SPARQL Query Processing Using Distributed Computation Engines

Mahmudul Hassan, Srividya Kona Bansal · 2018

The rapidly growing amount of linked open data demands semantic RDF services that are efficient, scalable, and distributed along with high availability and fault tolerance. To address this concern, the Big Data processing infrastructure Hadoop has been adopted for RDF data management systems. In this paper, we introduce distributed RDF data stores, namely VPExp and 3CStore, based on the existing vertical partitioning (VP) approach. In the VPExp approach, we propose splitting of predicates based on explicit type information of an object. The 3CStore scheme is designed with a 3-column store, comprising of a subset of triples from the VP table based on different join correlations, to reduce the number of join operations while executing SPARQL queries as SQL in a distributed system. We evaluate these two RDF data storage approaches by comparing them with vertical partitioning approach and state-of-the-art RDF management system S2RDF. We also present an evaluation of query performance of these systems built upon two popular distributed computation engines namely, Spark and Drill.

Read the paper · More papers on PaperTik