Mainstream Big Data Parallel Computing System Performance Optimization
Yarong Lv · Atlantis Highlights in Computer Sciences/Atlantis highlights in computer sciences · 2023
In recent years, with the widespread application of the Internet and information technology in people's production and life, the amount of data generated by all walks of life every day has shown a geometric and explosive growth trend, and the real-time nature of data analysis has become increasingly high.Therefore, big data parallel computing is widely used.The main purpose of this paper is to analyze and research the performance optimization of parallel computing systems based on mainstream big data.This paper mainly analyzes the design requirements and elastic resource scheduling strategy of the big data parallel computing system, introduces the framework and module design and the main functional modules, and preprocesses the data.The experimental results show that as the parallelism of Shuffle increases, the running time of Task decreases.This is because as the degree of parallelism increases, the amount of data processed by a single task decreases and the processing speed becomes faster.However, the proportion of shuffle data read I/O waiting time is the smallest only when the shuffle parallelism is 800, which is optimal.