EDP-ORD: Efficient distributed/parallel Optimal Rule Discovery

Sahar Mohamed Ghanem, Mona A. Mohamed, Magdy H. Nagi · 2011

Association rule discovery algorithms generate all rules satisfying minimum support and confidence thresholds. These techniques yield too many rules and are infeasible when the minimum support is low. Recently, Li proposed the Optimal Rule Discovery (ORD) algorithm that discovers a family of rule sets that maximizes a range of interestingness metrics, other than the commonly used confidence metric. In addition, the discovered optimal class association rule set is the minimum subset of rules with the same predictive power as the complete class association rule set. Moreover, ORD is significantly more efficient than association rule discovery independent of the data structure and the implementation. Due to the existence of huge amounts of data, it is important to investigate efficient methods for distributed/parallel mining of rules. In this paper, we propose EDP-ORD an efficient distributed/parallel extension of the ORD algorithm. We theoretically disclose a relationship between locally large and globally large rules and use it in reducing the number of generated rules and the exchanged messages at each site/partition. Moreover, we empirically compare EDP-ORD with a naïve distributed/parallel ORD version on five benchmark datasets. The experimental results shows that the reduction in number of generated rules at each site can reach 44% while the reduction in total size of exchanged messages can reach 58%.

Read the paper · More papers on PaperTik