Bucket MapReduce: Relieving the Disk I/O Intensity of Data-Intensive Applications in MapReduce Frameworks

Kai-Hsun Chen, Hsin-Yuan Chen, Chien‐Min Wang · 2021

Hadoop MapReduce is a software framework for processing vast amounts of data in parallel on large clusters. As data size increases, a need arises to resolve the significant increase in the disk I/O of reducer nodes, which may cause runtime bottlenecks. In response, we propose Bucket MapReduce, a system that utilizes bucket sort and pipeline parallelism to improve the performance of Hadoop MapReduce. Several parameters may influence the performance in Bucket MapReduce, and thus we will further discuss parameter tuning. In this paper, we perform experiments on TeraSort benchmark. Bucket MapReduce successfully reduces local disk I/O by 61% and improves the runtime by 1. 39× in 800GB TeraSort benchmark.

Read the paper · More papers on PaperTik