SFP: Processing of Small Files in Hadoop-based Distributed System

Seung-Hyun Kim, Young-Geun Kim, Wonjung Kim · 한국전자통신학회 학술대회지 · 2015

Mostly, mass data is processed in parallel in the distributed computing environment. Apache Hadoop is comprised of Hadoop Distributed File System (HDFS) and Map Reduce which support reliability, scalability, and distributed computing. However, Hadoop aimed at minimizing the number of data search times and transmitting the whole data set fast has its inherent limits. For the reason, it is not suitable to the analysis on a large data set consisting multiple small files. Hadoop Archive(HAR) or CombineFileInputFormat class provides flexible methods to deal with the issue. However, the methods are not a general solution applicable to all problems caused by small files. Therefore, to solve the problems, this study proposes a file packaging technique based on Small Files Package (SFP).

Read the paper · More papers on PaperTik