Improving Data Availability in HDFS through Replica Balancing

Rhauani Weber Aita Fazul, Paulo Vinícius Cardoso, Patrícia Pitthan Barcelos · 2019

Over time, the data distribution across an HDFS cluster may become unbalanced. The HDFS Balancer is a tool provided by Apache Hadoop that redistributes blocks by moving them from nodes with higher utilization to nodes with lower utilization. However, during block rearrangement, the HDFS Balancer does not aim to increase the availability of the data. This work presents a strategy that gives priority to block movements which increase the overall availability of the data stored in the HDFS. Thereby, increasing the fault tolerance as placing blocks in a higher number of racks tends to reduce the chances of data loss. In order to evaluate the implementation, an experimental investigation has been conducted to measure the system performance after balancing the cluster with the proposed solution.

Read the paper · More papers on PaperTik