Data Access Monitoring and Replication Control Management System for HDFS Clusters

Shantanu Shamraj, P. J. Kulkarni · 2018

Hadoop Distributed File System (HDFS) is very popular and widely used for storing large files across different nodes in the cluster. It has got many features and one of them is fault tolerance. Failure events are not avoidable. They may come at any point in time. Therefore, when client is working with business critical data which comes under hot data there should be some mechanism of fault tolerance to handle runtime node failures. In HDFS, 3-way replication strategy is used where two extra copies of data are present at other location in the cluster. However storage efficiency is very less in this case, as only 33% of the space can be used. In recent Hadoop Distribution, Erasure Coding scheme which is alternative to 3-way replication is introduced in which some parity blocks are calculated and stored in the cluster. This method ensures 67% storage efficiency. However it is not used for hot data in production clusters. It is generally applied to cold data. This paper focuses on monitoring data access and based on that labeling it either as hot or warm or cold and controlling its replicas in order to save space at runtime.

Read the paper · More papers on PaperTik