Heterogeneity-Based Files Placement in Hadoop Cluster
Suhas V. Ambade, Priya Deshpande · 2015
Hadoop which is an open source software framework uses MapReduce framework to process data. But its current implementation there is no direct support for file formats as well as it does not take heterogeneity amongst the cluster into consideration. With default block placement policy it distributes the data randomly. When there is heterogeneous environment so it's hard to utilise the power of each node correctly this may cause performance degradation. In proposed work we have added format specific data distribution amongst the nodes by using computational power of each node. So by experimentation, results shows there is improvement in the Map performance, MapReduce performance and overall performance of Hadoop in heterogeneous cluster.