An Algorithm of Data Skew in Spark Based on Partition
Shi Xiujin, Qian Yueqin · 2020
To solve the problem of data skew, many algorithms have been proposed at present. Due to different operating mechanisms, many advantages of hadoop-based algorithms cannot be fully realized in spark. However, most proposed algorithms are hadoop-based. Tang zhuo et al. proposed SKRSP, an adaptive partitioning method to deal with data skew in spark application. Compared with previous researches, this algorithm can more effectively alleviate the problems of data skew. Moreover, with the increase of data skew, the effect of this algorithm to deal with data skew is more and more significant. However, the research of this algorithm is based on the same hardware and software configuration of the nodes in the cluster. This paper presents a load balancing and key redistribution algorithm based on Spark (LBKRS) which optimizes the SKRSP algorithm from the point of view of load balancing. By monitoring the CPU utilization, memory utilization and other information of the calculation nodes, the LBKRS algorithm has a better effect on the data skew of different configuration nodes and is more adaptable to the actual production situation.