Distributed SPARQL query engine using MapReduce
Prasad Kulkarni · 2010
The Semantic Web is an emerging technology which aims at making data across the globe semantically connected. The data is represented in a very simple statement-like construct having a subject, predicate and an object. This can be visualized as a graph with the subject and the object as nodes and the predicate as an edge connecting the two nodes. When many statements like these are collected together they forms an RDF graph. There are RDF query languages to query such data, and SPARQL is one of them. According to the SP2Bench performance benchmarks, the SPARQL queries are very slow for RDF data with millions of triples. Hence, we aim to develop a distributed SPARQL query engine using the MapReduce (introduced by Google) model of parallelization and hypothesize that this system will outperform the scalability and performance reported by the SP2Bench. We extend ARQ, an open source SPARQL query engine provided by the Jena framework, to work with the Hadoop MapReduce framework and implement distributed SPARQL query processing. This thesis provides the detailed implementation and algorithmic details of our work. We contribute two novel methods to optimise RDF query engine which exploits document indexes and a join pre-processing technique. The experimental results show the merits and demerits of using MapReduce for distributed RDF query processing and provides us a clear path for future work.