Minimizing Skew in MapReduce Applications Using Node Clustering in Heterogeneous Environment
Vishal Ankush Nawale, Priya Deshpande · 2015
We present an automatic skew minimization approach defined for MapReduce programs and present proposed system that implements this approach as a replacement for an existing MapReduce implementation. The proposed system addresses these challenges and works as follows: From intermediate output from map tasks data skew present in records to solve this problem we create two sets of node mainly of high and low computational power nodes and assign the skewed records to the high computational power nodes and remaining to the low computational power nodes to process further reduce task. We implement proposed system as an extension to Hadoop and evaluate its effectiveness using real applications. The results show that proposed system can reduce job runtime in the presence of skew and adds little to no overhead in the absence of skew in heterogeneous environment.