Research on Optimization Strategy of Data Skew Problem Based on Hive Partitioning Mechanism

Hongyuan Wang, Xingdong Ye · 2023

Due to the complex structure and large scale of big data, there are also data skew problems caused by uneven data distribution and the limitations of default partitioning algorithm, which is also a common performance bottleneck in distributed computing, and seriously affects the overall resource utilization of the system. Based on the execution principle of MapReduce, this paper proposes Hive based load balancing methods such as random prefix, preaggregation, parameter tuning, etc., in order to improve computing parallelism and reduce application completion time.

Read the paper · More papers on PaperTik