Strategy of storing and accessing small web log files on Hadoop
Qiucheng Ban, Zhengping Jin · 2017
Hadoop distributed file system(HDFS) was designed to manage large files. For a large number of small files, such as genomic data and web log data, HDFS would suffer from low access and storage efficiency and bring much less benefit. In this paper, we propose an asynchronous merging algorithm based on monitoring the task queue (MTQM) and a prefetching strategy based on Hash index. On the basis of client merging small files, uploading and merging small files can be asynchronous by monitoring the task queue. Meanwhile, small files can be quickly accessed by using prefetching and caching scheme. The experimental results show that the approach proposed in this paper can improve the efficiency of storing and accessing small files and effectively reduce the metadata storage space of NameNode.