Evaluation of Performance Saturation Using the Hadoop Framework

Rafael Sobrinho Ferreira, Bruno Guazzelli Batista, Rafael de Magalhães Dias Frinhani, Bruno Tardiole Kuehne, Dionisio Machado Leite Filho, Maycon Leone Maciel Peixoto · 2018

It is estimated that about 2.5 exabytes of data are produced daily. This large volume of data has brought new possibilities of applications, however, to manage this large volume of data, new technologies were needed. One of the most prominent technologies is the Hadoop framework, which implements a parallel task processing paradigm. The aim of this paper is to present the results of our group's research which analyzed the performance of the Hadoop framework for Big Data processing. The performance evaluation focused on finding the saturation point of Hadoop performance by varying the number of nodes in the cluster applying two benchmarks - TeraSort and Pi. The analysis was performed using a real infrastructure, implementing the system in a physical cluster, providing a general approach of performance analysis in the Hadoop framework for developers and researchers.

Read the paper · More papers on PaperTik