Load Balancing Large Data Sets in a Hadoop Cluster
Andriavelonera Anselme A, Rivosoaniaina Alain N, Rakotomalala Francis, Thomas Mahatody, Manantsoa Victor · International Journal of Distributed and Parallel systems · 2023
With the interconnection of one to many computers online, the data shared by users is multiplying daily. As a result, the amount of data to be processed by dedicated servers rises very quickly. However, the instantaneous increase in the volume of data to be processed by the server comes up against latency during processing. This requires a model to manage the distribution of tasks across several machines. This article presents a study of load balancing for large data sets on a cluster of Hadoop nodes. In this paper, we use Mapreduce to implement parallel programming and Yarn to monitor task execution and submission in a node cluster.