Load balancing greedy algorithm for reduce on Hadoop platform

Hui Xia · 2018

MapReduce is a widely used parallel computing framework and is an important part of the Hadoop platform. It mainly includes Map and Reduce functions, and Map function outputs key-value key-value pairs as input for Reduce. Due to the dynamic input, Reduce How to solve the load balancing of Reduce is an important research direction to optimize MapReduce. Taking the whole data as a sample and analyzing the data by the appropriate amount of samples to get a reliable key distribution with a small cost, Algorithm instead of the default Hash algorithm of Hadoop platform to divide the data into Reduce load balancing. The main idea of the proposed greedy algorithm is to calculate the average of all key frequencies and the number of Reduce nodes according to the sampling data and then allocate each Reduce A load that is close to the average, so as to achieve the overall load balancing. The simulation results show that compared with the default hash partitioning algorithm, the proposed algorithm can save 10.6% uptime and achieve better load balancing.

Read the paper · More papers on PaperTik