Concurrent Spark message distributor

Yulin HE, Zejie LIN, Yuanyuan Xu, Yingchao Cheng, Zhexue HUANG · JOURNAL OF SHENZHEN UNIVERSITY SCIENCE AND ENGINEERING · 2025

In the Spark big data computing framework, the driver employs an iterative message distributor mechanism that incurs considerable task submission overhead, delays task initiation, restricts execution concurrency, and causes idle waiting among executors—ultimately leading to inefficient utilization of computing resources. To address these issues, we propose an efficient and lightweight concurrent Spark message distributor based on a thread pool scheduling strategy. In contrast to Spark's original distributor, the proposed design is better suitable for scheduling fine-grained, high overhead tasks. It parses metadata containing key executor information to extract the task list and corresponding executor identifiers to each task, then initializes a thread pool to launch asynchronous computations for each task, thereby enabling true concurrent task distribution. This approach significantly reduces dispatch latency while ensuring system stability and reliable task execution. Experimental evaluations conducted in a virtualized cluster environment demonstrate the superiority of the proposed distributor over the original Spark mechanism. Results show that, with memory usage held constant, the concurrent distributor reduces task execution time by about 9% and increases central processing unit utilization by about 5%. The proposed concurrent Spark message distributor, effectively mitigates the high overhead and computational resource inefficiency associated with traditional message distribution methods in fine-grained task scenarios.

Read the paper · More papers on PaperTik