YARM:Efficient and Scalable Semantic Reasoning Engine Based on MapReduce
GU Ron · Chinese Journal of Computers · 2015
The rapid development of the Semantic Web has produced massive amount of the RDF data.The major challenge for large scale RDF semantic reasoning is that it involves huge amount of computation.This makes the whole process very time-consuming.It is obvious that the traditional semantic reasoning engines are not efficient when dealing with the massive amount of RDF data.On the other hand,the state-of-art distributed semantic reasoning algorithms built with MapReduce lack of optimization for reasoning process in a distributed and parallelized environment.Thus,this still makes the reasoning process relatively time-consuming.In addition,most of existing reasoning engines lack of scalability.To solve these problems,we design and implement YARM,a new parallel semantic reasoning algorithm and engine that built with the MapReduce parallel model.YARM includes four major optimizations:first,it adopts a welldesigned data partitioning schema and a corresponding reasoning algorithm to minimize the amount of data transferred among computing nodes;second,it optimizes the execution order ofthe reasoning rules to improve the computing speed;third,it uses an efficient way to remove duplicates yielded in reasoning process.This avoids the need of extra MapReduce jobs to do this work;forth,based on the optimizations above,we design and implement a new parallel reasoning algorithm on the Hadoop MapReduce framework.Experimental results on both real-world and synthetic datasets show that YARM is about 10 times faster than the latest MapReduce-based reasoning engine and also achieves better scalability.