Intermediate Data Placement Strategy for Different Data Skew Levels Based on Random Sampling in Spark

Xueqian Gong, Chunlin Li, Youlong Luo · 2019

In recent years, the Apache Spark had been widely used in processing large-scale data. However, when the input data onto MapReduce was skewed, the default intermediate data placement algorithm of the Apache Spark could not efficiently handle the skewed data. In this paper, in order to achieve load balancing with all reducers, the focus on our attentions was how to process data efficiently when the input data onto MapReduce was skewed. So we defined a data skew model to measure the skew degree of the intermediate data. We also applied a sampling algorithm based on reservoir model. When the skew of input data was severe, we proposed a fine-grained data placement algorithm based on splitting cluster. When the skew of input data was slight, we proposed a coarse-grained data placement algorithm that did not split the clusters. In our experiment, we first chose the appropriate sampling rate. Then, we determined the optimal value of the parameter that measures the two degrees of the skewed input data. At last, we compared the average execution time of several algorithms and the load balancing degree of reducers under different conditions. The experiments confirmed the efficiency of the two intermediate data placement algorithms proposed in this paper in their respective usage scenarios.

Read the paper · More papers on PaperTik