The main characteristics of five distributed file systems required for big data: A comparative study

Akram Elomari, Larbi Hassouni, Abderrahim Maizate · Advances in Science Technology and Engineering Systems Journal · 2017

These last years, the amount of data generated by information systems has exploded.It is not only the quantities of information that are now estimated in Exabyte, but also the variety of these data which is more and more structurally heterogeneous and the velocity of generation of these data which can be compared in many cases to endless flows.Now days, Big Data science offers many opportunities to analyze and explore these quantities of data.Therefore, we can collect and parse data, make many distributed operations, aggregate results, make reports and synthesis.To allow all these operations, Big Data Science relies on the use of "Distributed File Systems (DFS)" technologies to store data more efficiently.Distributed File Systems were designed to address a set of technological challenges like consistency and availability of data, scalability of environments, competitive access to data or even more the cost of their maintenance and extension.In this paper, we attempt to highlight some of these systems.Some are proprietary such as Google GFS and IBM GPFS, and others are open source such as HDFS, Blobseer and AFS.Our goal is to make a comparative analysis of the main technological bricks that often form the backbone of any DFS system.

Read the paper · More papers on PaperTik