Privacy Preserving Rack-Based Dynamic Workload Balancing for Hadoop MapReduce
Xiaofei Hou, Doyel Pal, Ashwin Kumar T.K., Johnson P. Thomas, Hong Liu · 2016
Hadoop has two components namely HDFS and MapReduce. Hadoop stores user data based on space utilization of datanodes on the cluster rather than the processing capability of the datanodes. Furthermore Hadoop runs in a heterogeneous environment as all datanodes may not be homogeneous. For these reasons, workload imbalances will occur when jobs run in a Hadoop cluster resulting in poor performance. In this paper, we propose a dynamic algorithm to balance the workload between different racks on a Hadoop cluster based on the log files of Hadoop. However, if the tasks are executing on critical or sensitive data in a secured rack, the data transfer to an unsecured cluster will result in privacy being compromised. We propose an approach to transfer data between racks without disclosing private information. Moving tasks from the most overloaded rack to another rack improves the performance of MapReduce jobs. Our simulations indicate that the proposed algorithm decreases running time of a job by more than 50% running on the most overload rack.