Performance efficiency in Hadoop for storing and accessing small files

Awais Mehmood, Muhammad Usman, Waqas Mehmood, Yasmeen Khaliq · 2017

Hadoop Distributed File System (HDFS) is a famous distributed file system to access and store large files. It suffers from different memory, page fault, and processing bottlenecks while accessing storing a large number of very small sized files. In this research, we have studied architecture, operation, and behavior of HDFS while serving the access and storage requests for a very large number of very small-sized files. Hadoop provides a solution to this problem i.e. archiving these many small files with a single HAR file and treating this HAR file as a large file. In this study, we have analyzed the working and significance of HAR as well as its cost benefit analysis is done in a comprehensive manner.

Read the paper · More papers on PaperTik