A dynamic data placement policy for heterogeneous Hadoop cluster

Santa Maria Shithil, Tushar Kanti Saha, Tanusree Sharma · 2017

Hadoop Distributed File System (HDFS) is a file system for storing and managing big data. The current HDFS block placement policy works well for the homogeneous cluster. However, it can not evenly and fairly distribute blocks across the heterogeneous cluster and results in an unbalanced cluster. An unbalanced cluster also reduces MapReduce performance. Hadoop relies on load balancer tool to balance replica distributions that in turn degrade the overall performance of hadoop. In this paper, we proposed a dynamic block placement policy that distributes blocks across a heterogeneous cluster more evenly than the current HDFS block placement policy and also improves the performance of MapReduce applications. The solution had been implemented and test had been conducted to evaluate its contribution to Hadoop.

Read the paper · More papers on PaperTik