SFSAN Approach for Solving the Problem of Small Files in Hadoop
Tharwat El-Sayed, Mohammed Badawy, Ayman El‐Sayed · 2018
Hadoop is a distributed computing framework written in Java and used to deal with big data; it is designed to handle large files. Handling the small files leads to some problems in Hadoop performance. Several approaches are used to overcome the housing and the access efficiency problems of the small files in Hadoop. These approaches are Hadoop Archive Files, Federated NameNodes, Batch file consolidation, Sequence files, HBase, and S3DistCp. All these approaches have some advantages, but they have some limitations as the slow in converting all the HDFS Files to Sequence Files and the lack of correlations between the files stored in the Sequence Files. In this paper, we proposed an enhancement of the Sequence files approach called Small Files Search and Aggregation Node (SFSAN) approach. Our proposed approach improves the Hadoop performance by overcoming some of the limitations of the Sequence Files approach and keeping up its advantages.