Algorithm Rebuilding and Performance Optimization of MapReduce in Hadoop
Ben Y. Zhao · Institutional Repositories DataBase (IRDB) · 2014
Abstract-—As a core component of Hadoop that is a cloud open platform, MapReduce is a distributed and parallel computing model based on mapping function for processing and generating large data sets. MapReduce abstracts business logic from implementation details, and provides powerful interfaces for programmers to use. It can mask the underlying speciflc implementation processes and efficiently reduce the distributed and parallel computing difficulty, and has high reliability, high scalability, high efficiency, high fault tolerance features. However, MapReduce mechanism itself is not perfect and mature, and needs to be improved the efficiency further. According to analyzing mapreduce principles and performance indicators, in heterogeneous environments, unreasonable resource scheduling, data transmission and system parameters are summarized. By rebuilding algorithms and optimizing performances, sliding window scheduling algorithm (SWSA), changing transfer protocol from HTTP to UDT and optimization of system configuration parameters arc proposed. At last, this paper compares performance difference of MapReduce before optimization with after optimization. The algorithms arc verified by experiments, rebuilding and optimization greatly improve performance of hadoop.