Handling Skewed Data in Mapreduce Environment

Ankit Kumar · 2015

The MapReduce framework is a parallel processing paradigm which uses shared nothing architecture for processing big data on a distributed cluster. However, MapReduce systems have become well known for processing large data sets and are more significantly being used in scientific applications. In contrast to simple application scenarios, scientific applications involve complex computations which poses challenge to MapReduce systems. Particularly because, (a) processing time complexity of the mapreduce task is generally high and (b) scientific data is frequently skewed. When the data is heavily skewed, the data affects the performance of operations in parallel environment (such as mapreduce) where data is distributed among parallel nodes for processing. Data skew can greatly limit the efficiency of parallel processing when some processing nodes are overloaded during data distribution and hence take a greater time for completion as compared to other processing nodes. This also results in wastage of resources of the idle processing nodes. As data skew naturally occurs in many applications, handling it is an important issue for improving the performance of the join operation. We implemented a skew partitioning algorithm which has the ability to handle skew and conducted rigorous experiments, to prove that our method gives efficient results when compared to the other optimization methods.

Read the paper · More papers on PaperTik