Optimization of a Network Retrograde Analysis System Implemented on HDFS
Xiaofang Yuan, Wubin Zhang · 2019
The Apache Hadoop is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. In addition, it provides a distributed file system (HDFS) that stores data on the compute nodes, providing very high aggregate bandwidth across the cluster. However, the problem of small-file storage in Big data is restricted by the ability of HDFS to work effectively, as it is originally designed to deal with large files. In order to construct a network retrograde analysis system (RAS), we propose a method to optimize the performance of HDFS for small file storage. In this system, a large number of network packet (pcap) files are stored in HDFS, each 40 to 50 MB in size. To reduce the number of files, pcap files are merged into larger files, meanwhile indexes for each larger file are also generated. A distributed caching index strategy is introduced to improve the read speed of pcap files in HDFS. The experimental results show that our system can improve the access efficiencies of small files in HDFS, while reducing the NameNode memory consumption by about 90 compared to the original HDFS, and 70% to Hadoop Archive program (HAR), respectively. The access time per MB file is reduced remarkly, by about 50% lower than the original HDFS, and 70% than that of HAR.