A Spark Join Algorithm Based on Memory Monitoring and Batch Processing
Kefei Cheng, Zhao Luo, Zhou Ke, Xianjun Deng, Xudong Chen · 2018
In recent years, the Spark memory computing framework has risen rapidly, and the data processing speed has been greatly improved. However, the upper limit of speed is limited by the Spark memory size. The most typical one is the large table Join algorithm, which uses the Sort Merge Join algorithm by default. For the problems such as frequent memory shortage (OOM), serious GC, large Shuffle overhead, and heavy time consuming, we propose a Join algorithm based on memory monitoring and batch processing, which is based on the signaling data scene of LTE network. We first analyze four aspects: data files, memory usage, batch processing, and runtime monitoring. Then, in the data organization mode of Parquet+Snappy, we control the data flow of Join in batches through memory monitoring, and optimize the default Join by combining BloomFilter and local Hash Join. The experiments show that our optimized Join algorithm alleviates the problem of insufficient memory to some extent and reduces the load imbalance caused by data skew. The overall running time is better than that of the default Join algorithm.