Improving the MPI Remote Memory Access Model for Distributed-memory Systems by Implementing One-sided Broadcast

Mohamed M. Abuelsoud, Alexey A. Paznikov · 2024

Currently, processing large volumes of expanding data efficiently and consistently is a significant challenge. Traditional distributed-memory high-performance computers (HPC) based on message-passing model struggle with inherent synchronization difficulties, limiting their ability to keep pace. Remote Memory Access (RMA, also known as one-sided MPI communications) allows a process to directly read from or write to the memory of another process, bypassing the need for message exchange. Unfortunately, there is no collective operation interface in the current MPI RMA standard. However, RMA has the potential to reduce synchronization costs by enabling concurrent access to shared data structures, distributed among MPI processes’ memories. Existing onesided MPI standards offer a linear interface only that hampers parallelization and far from efficient. To bridge this gap, we propose an algorithm design for efficient collective (parallelizable) operations in the RMA paradigm. Our study primarily examines the benefits of collective operations using the broadcast algorithm as an example. Our implementations surpass traditional methods, demonstrating the promising potential of this technique, as more performance tests indicate.

Read the paper · More papers on PaperTik