Improving the performance of heterogeneous Hadoop cluster

Ch. Bhaskar Vishnu Vardhan, Pallav Kumar Baruah · 2016

With the development of information Technology, we can see an exponential growth in the amount of data that is being generated and used over the last decade. The need for storing the data and processing them in order to extract value from skew data has paved way for parallel, distributed processing applications like Hadoop. The current Hadoop implementation assumes that the compute capacity of all the nodes in a cluster are homogeneous in nature, but looking at the cloud infrastructure, we can see that different hardware configurations systems are being used which is logical. Hence, it is necessary to study the data placement policy where we can distribute the data based on the processing power of a node. In this paper, we propose a dynamic block placement strategy in Hadoop to distribute the input data blocks to the nodes based on the computing capacity of each node. The proposed algorithm which can adapt and balance data dynamically and reorganize the input data in HDFS according to the compute capability of each node in hadoop heterogeneous enviornment. The proposed method reduces the data transfer time to achieve improved performance. We ran benchmark applications against the proposed block placement strategy and original HDFS blocks placement strategy. The experimental results show that the dynamic data placement strategy can decrease the execution time and improve performance of hadoop heterogeneous cluster. Hadoop is designed to handle Big files and with less frequent updates. But nowadays many applications need to handle small files, which has become a pertinent issue for degradation of performance of an application on hadoop platform. Our proposed methods have shown good improvement of performance in these cases. we are able to see a speedup of 25.7x and 20x compared to the original Hadoop when ran on wordcount and grep benchmark applications.

Read the paper · More papers on PaperTik