When Rule Engine Meets Big Data: Design and Implementation of a Distributed Rule Engine Using Spark
Jindou Zhang, Jinxing Yang, Jing Li · 2017
Rule Engines have been widely used both in industry and academia since they can separate rule knowledge from implementation logic conveniently and flexibly. However, traditional rule engine systems can not deal with big data, because of the limitations of memory and computing capacity of one single computer. Consequently, some researchers have proposed distributed rule engine to meet this challenge. But these solutions still do not work well when enormous amounts of facts are involved, for the reasons such as high cost of moving data, data imbalance and so on. We present the SparkRE system, a distributed rule engine based on Spark, to support big data reasoning. In particular, SparkRE uses DataFrame representing distributed working memory which holds enormous amounts of facts, and implements the rule condition-testing mechanism by the query execution engine of Spark SQL. In addition, we design a rule language tailored for SparkRE and an efficient inference engine. An experimental evaluation shows SparkRE's capability of matching 26 million facts, and reveals its scalability and performance.