Assessing and improving the performance and scalability of an iterative algorithm for Hadoop

Joao Paulo Barbosa Nascimento, Daniel de Oliveira Capanema, Adriano C. M. Pereira · 2017 Computing Conference · 2017

Large volumes of data are generated and collected nowadays through sensors, devices, social networks, among others. The ability to handle large datasets has become an important expertise for the success of many organizations around the world, demanding increasingly parallel and distributed processing. Parallel computing, based on multicore architecture, re-emerged in recent years as a solution for this challenge. To help developers to design parallel programs, there exists several parallel and distributed tools (frameworks), such as Apache Hadoop and Spark. These frameworks provide a plenty of configuration parameters (e.g., Hadoop has more than 200) and to configure all of them optimally is a non-trivial task. This work investigates the influence of various parameters on the performance of Apache Hadoop when process HEDA, an iterative algorithm that calculate metrics of centrality in large graphs. The execution of HEDA in a complex network is extremely important because there are several centrality measures that determine the importance of a node in the graph. The centrality of a node is a key issue in the analysis of complex networks, especially on social networks. It was observed that in some cases the improvement in execution time reached almost 80% by applying the proposed values in the selected parameters. Moreover, it was possible to increase CPU consumption four times and to achieve a Speedup of 92% of linearity. The work also describes the steps that were followed to perform the experiments, which can help other researchers to conduct new experiments.

Read the paper · More papers on PaperTik