Locality-based Partitioning for Spark

Yuchong Xia, Fangfang Yang · 2017

Spark is a memory-based distributed data processing framework.Lots of data is transmitted through the network in the shuffle process, which is the main bottleneck of the Spark.Because the partitions are unbalanced in different nodes , the Reduce task input are unbalanced.In order to solve this problem, a partition policy based on task local level is designed to balance the task input.Finally, the optimization mechanism is verified by experiments, which can alleviate the data-skew and improve the efficiency of the job process.

Read the paper · More papers on PaperTik