Breaking the Communication Bottleneck: A Novel MPI Framework Based on Priority Flow Scheduling for Distributed Scientific Computing Tasks

Jiangping Han, J. B. Liu, Yi Liu, Yu Shen, David S. L. Wei, Kaiping Xue · IEEE Communications Magazine · 2025

Scientific computing tasks are often distributed across multiple servers for parallel computing using message-passing interface (MPI) libraries. However, inter-node communication costs have become a significant bottleneck, adversely affecting task completion times. This article analyzes the characteristics and traffic patterns of scientific computing tasks, identifying the urgency and transmission needs of different communication flows. Based on these findings, we propose PFS, a novel MPI framework that employs priority flow scheduling. PFS assigns fine-grained priorities to communication flows generated by individual computing tasks via a cross-layer design. It uses quality of service (QoS) queues for preemptive scheduling, providing differentiated service to various types of traffic with varying transmission requirements. Implemented in OpenMPI and tested with OpenFOAM at a supercomputing center, PFS reduces communication time between nodes and decreases task completion time by up to 10.15% compared to benchmarks.

Read the paper · More papers on PaperTik