A Review of Various Optimization Schemes of Small Files Storage on Hadoop

Liang Huang, Jun Liu, Meng Weixiu · 2018

With the rapid development of electronic information industry and the increasing amount of Internet traffic, the amount of data generated grows exponentially. Such a large scale of data has gone far beyond the capacity of traditional storage devices. Owing to its powerful data storage and distributed computing abilities, Hadoop has become very popular big data distributed processing software framework in recent years, and performs well when storing and managing super big data sets. However, the performance of HDFS would also be suffered when handling a large number of small files since they place a heavy burden on the NameNode of HDFS both in terms of memory and access time. Therefore various solutions have been proposed by experts and scholars to overcome these defects. In this paper, the basic architecture of Hadoop system is firstly introduced, and problems, which are generated when Hadoop handles a large number of small files, are analyzed and summarized, and the necessity of optimization scheme of small file storage based on Hadoop is showed. Then, some existing optimization methods are introduced in detail, and the ideas, advantages and disadvantages of methods of relevant literature are analyzed briefly. Finally, combined with the actual situation of the small file storage based on Hadoop, some opinions and suggestions are proposed.

Read the paper · More papers on PaperTik