SF-SNF: A Small-file Sniffer System for Hadoop Clusters
Miao Wang, Yanming Zhang, Hualei Wang · 2024
Apache Hadoop has been a major component in the big data ecosystem for more than a decade. It relies on the Hadoop Distributed File System (HDFS) to store large datasets and MapReduce to process these distributed datasets. HDFS manages the metadata of all its files through a server known as the Namenode. To achieve high availability (HA), Hadoop clusters typically deploy two Namenodes: one active and one standby. This architecture enables Hadoop to store and process massive files with good reliability. However, HDFS often encounters significant performance degradation when managing a large number of small files. Scanning all files to locate the small ones through the Namenode’s service becomes time-consuming and adds extra burdens to the Namenodes. There is a lack of research on how to identify the hotspots of small files in HDFS without querying the Namenode in a time-efficient manner. In this paper, we designed a big data system to identify small files in HDFS by parsing the File System Image (FSImage), which is generated periodically on the Namenode. This system utilizes the standby Namenode to parse the FSImage and send the file information to an Apache Kafka topic. A Doris Routine Load procedure then listens to the topic and loads the information into a table containing the file information for real-time querying by users. This approach allows Hadoop cluster administrators to identify small files using the standby Namenode without impacting the normal operation of the Hadoop cluster. Additionally, it can serve as a tool to locate small-file hotspot tables in a Hive data warehouse.