Distributed RDFS Rules Reasoning for Large-Scaled RDF Graphs Using Spark

Ren Li, Qi Zhang, Huibin Wang, Guiping Wang · 2016

Scalable processing on large-scaled RDF graphs becomes a critical issue with the explosion of semantic web technologies. Most of the existing distributed RDF querying and reasoning solutions are designed based on the MapReduce paradigm. However, MapReduce should be further optimized since several inherent limitations such as lack of efficient job scheduling and iterative computing mechanisms affect its performance and flexibility. To overcome the drawbacks, some novel distributed programming models like Spark have been released and comprehensively used. To further improve the efficiency of RDFS rules reasoning for large-scaled RDF data, this paper design a graph-based RDF data partitioning and storage schema based on HBase. A novel RDFS reasoning approach is proposed by exploiting the Spark context. An experiment on the standard LUBM benchmark shows that our approach is more efficiency than existing solution.

Read the paper · More papers on PaperTik