Node Failure Recovery Algorithm for Distributed File System Based on Measurement of Data Availability

Xingyao Yang · 2013

The strategy for distributed file system dealing with node failure needs much bandwidth and disk space resources and affects stability of the system.By studying HDFS's cluster structure,data blocks storage mechanism,the state relationship between node and block,we defined the cluster nodes matrix,node status matrix,file block partition matrix,block storage matrix and block state matrix.Those definitions enable us to model the availability of data block easily.Based on the measurement of data block's availability,we proposed the new node failure recovery algorithm and analyzed the performance of the algorithm.The experimental results show that compared with the original strategy,the new algorithm ensures the availability of all blocks in the system and reduces the bandwidth and disk space resources for recovery,shorts the recovery time,and improvs the stability of system.

Read the paper · More papers on PaperTik