Optimizing Hadoop Distributed File System Replication Policies with Predictive Categorization
Nada A. Zayed, Yasmine N. M. Saleh, Ahmed A. Aboelfarag, Mohamed Shaheen · 2024
The rapid increase in data volume over the past years has given rise to big data science. Distributed File Systems (DFS) have become widely employed to handle this vast amount of data, such as Google File System and Hadoop Distributed File System (HDFS). The primary objective of a DFS is to ensure data availability and system reliability in the event of failure. Data availability and system reliability are achieved by replicating files across multiple locations, which, however, results in the consumption of storage space and other resources. The significance of these files varies based on their frequency of use within the system. As a result, specific files are deemed less critical and do not warrant extensive replication, as they hold little importance in the overall system. This paper presents a novel approach called "Optimizing HDFS Replication Policies with Predictive Categorization" for storage efficiency. This approach aims to minimize storage consumption while ensuring data availability and system reliability. Experimental results using the Spotify Hit Predictor Dataset (1960-2019) showcase significant improvements in storage utilization and throughput. This approach not only enhances system performance but also positions organizations for substantial cost efficiencies in managing their data-intensive workloads.