Parallel Processing SPARQL Theta Join on Large Scale RDF Graphs

Tao Wang, Pingpeng Yuan, Xiaofei Liao, Hai Jin · 2018

Theta join is commonly employed in real practices. Although SPARQL is a popular RDF query language, SPARQL does not define the specification of theta join until 2013. After that, few RDF stores except the stores based on RDBMS or key-value stores can process theta join. However, processing theta join on RDBMS or key-value stores is costly due to lack of RDF-native optimization. Efficient solutions to process theta join queries on RDF graph directly have not been fully explored. Here, we present ThetaStore to parallel process SPARQL theta join queries on large RDF graphs. First, a uniform SPARQL query graph model for theta join queries and equi-join queries is defined. Second, each subject, predicate, and object of RDF triples are mapped into order-preserving IDs instead of an random integer as most of RDF stores do. By this way, the engine does not need to translate theta join into a set of equi-joins. Finally, the engine employs decomposition and segment-oriented parallelization to execute SPARQL queries. The engine decomposes query into several star sub-queries and the intermediate results matching each pattern are divided into segments. Extensive experiments are conducted on large RDF data sets and experimental results show that ThetaStore outperforms state-of-the-art systems in both theta joins and equi-joins.

Read the paper · More papers on PaperTik