Merging and cooperative caching based HDFS small files access performance optimization

Weijing Xu, Yitao Qiu, Congfeng Jiang, Junming Liu, Xiaoyu Wang, Shi‐Jie Chen · 2024

In this paper, we propose a performance optimization method for accessing small files in HDFS based on merging an d co-caching strategies to address the memory overhead problem caused by massive small files in Hadoop Distributed File System (HDFS). First, the small files stored in HDFS are merged into multiple large files based on the descending first adaptation algorithm, thus saving the memory space occupied by the management node (NameNode). Second, in order to improve the read speed of s mall files, this paper proposes a two-level collaborative caching strategy, i.e., the server-side caching adopts the improved GDSF (G reedy Dual Size Frequency) strategy for small files, and the client -side adopts the ARC (Adaptive Replacement Cache) caching strategy for the small file indexes. The Experiments show that compared with genetic algorithm and particle swarm algorithm, the number of large files generated by merging small files using descending first adaptation algorithm is reduced by 26 and 60 on average. Compared with the original GDSF strategy, the co-caching strategy proposed in this paper improves the server-side cache hit rate by 8.2% on average. Compared with the FIFO and LRU cache re placement strategies, the client-side cache hit ratio is improved by 2.7% and 8.9%, respectively.

Read the paper · More papers on PaperTik