Design of an Optimal Data Placement Strategy in Hadoop Environment
Shah Dhairya Vipulkumar, Saket J Swarndeep · International journal of advance research and innovative ideas in education · 2017
The MapReduce framework has gained wide popularity as a scalable distributed system environment for efficient processing of large scale data of the order of Terabytes or more. Hadoop, an open source implementation of MapReduce coupled with Hadoop Distributed File System, is widely applied to support cluster computing jobs requiring low response time. The current Hadoop implementation assumes that nodes in the cluster are homogenous in nature. Data placement and locality has not been taken into account for launching speculative processing tasks. Furthermore, every node in the cluster is assumed to have same CPU and memory capacity despite some of the nodes being configured using vastly varying generation of hardware. Unfortunately, both the homogeneity and data placement assumptions in Hadoop are optimistic at best and unachievable at worst, potentially introducing performance problems in Hadoop clusters at data centres. This dissertation explores the Hadoop data placement policy in detail and proposes a modified data placement approach that increases the performance of the overall system. Also, the idea of placing data across the cluster according to the processing capacity utilization of the nodes is presented, which will improve the workload processing in Hadoop environment. This is expected to reduce the response times for the applications in large-scale Hadoop clusters.