Optimization of hadoop small file storage using priority model
V. Nivedita, J. Geetha · 2017
The improvement in the technology and the desperate need to store huge data has been increasing steadily. Hadoop uses MapReduce framework for processing big datasets, which consists of namenode and datanode. The hadoop framework technology has been extensively used to store, process, and use the enormous data that are being placed into server. Handling of huge small files is becoming laborious since the namenode has to manage the filenames and its corresponding metadata. For the purpose of fault tolerance, the data needs to be replicated on the data nodes. In order to handle the mammoth number of small files and to reduce the out of memory burden on the namenode, several techniques are being approached which involves HDFS, EHDFS, SFS, HAR, NHAR. In this paper we discuss about the different techniques that are developed to handle the storage of small files and propose a new method to solve the storage issue in hadoop. The new method takes the priority of the file as a distinguishing factor which further helps in reducing the memory usage of namenode.