A Clustering-Aware and Multi-Feature Replication Strategy for Efficient Storage in HDFS

Helmi Tlich, Hanène Chettaoui, Tarek Hamrouni · Procedia Computer Science · 2025

Default replication strategy in Hadoop Distributed File System (HDFS) comes with a fixed replication factor that does not take the file’s access patterns or usage into account. This static approach leads to unnecessary storage overhead, as cold data gets over-provisioned, while hot data gets under-provisioned. In this paper, we propose a novel replication approach based on unsupervised clustering using K-Means++. In order to generate a normalized feature space, our solution extracts five core features from HDFS logs, namely access frequency, file age, write ratio, locality, and concurrency. K-Means++ is then used to uncover natural groupings of files with similar usage patterns. Each cluster is later matched with a replication strategy tailored to its needs: critical files receive additional replicas, shared files benefit from improved locality through smarter placement, and archival data is stored more efficiently using erasure coding. The evaluation of the implemented framework prototype over a virtualized HDFS environment shows that our approach is capable of increasing overall storage savings by over 23%, improving the read latency of hot data by up to 30%, and reducing network bandwidth consumption by 20%, all while maintaining fault tolerance.

Read the paper · More papers on PaperTik