Efficien t Processing Distributed Joins with Bloomfilter using MapReduce y
Changchun Zhang, Lei Wu, Jing Li · 2013
The MapReduce framework has been widely used to process and analyze large-scale datasets over large clusters. As an essential problem, join operation among large clusters attracts more and more attention in recent years due to the utilization of MapReduce. Many strategies have been proposed to improve the efficiency of dis-tributed join, among which bloomfilter is a successful one. However, the bloomfilter’s potential has not yet been fully exploited, especially in the MapReduce environmen-t. In this paper, three strategies are presented to build the bloomfilter for the large datasets using MapReduce. Based on these strategies, we design two algorithms for two-way join and one algorithm for multi-way join. The experimental results show that our algorithms can significantly improve the efficiency of current join algorithm. Moreover, cost models of these algorithms are characterized in order to find out the way of improving the performance of two-way and multi-way joins.