Enhancing Storage Efficiency of HDFS for Small Files by Dynamic File Pooling
Raveena Aggarwal, Jyoti Verma, Manvi Siwach · 2024
The technology is evolving fast. The data is growing faster. Apache Hadoop comes up as a capable platform to foster this rapidly progressing data. However, it is quite evident that the framework for Hadoop has been designed to manage large files. As soon as it has to face a lot of small files, the performance sinks. This further leads to metadata overheads, processing overheads, network clogging and a list of other issues. This paper proposes a novel framework for merging these small files into large files known as pools. These pools can be updated dynamically whenever subsequent datasets are fed into the system according to the empty space available. Dynamic file pooling significantly improves the storage efficiency of Hadoop Distributed File System (HDFS). The results are obtained over a smaller dataset which can be scaled for over a million of files in real scenario.