Evaluating One-Sided Communication on Graph500 with MPI-RMA and OpenSHMEM

Jefferson Boothe, Alan D. George · 2024

While traditionally utilized for two-sided and collective communication, the latest MPI standards support remote memory access (RMA) between processes to enable one-sided communication. This paradigm is typically associated with fine-grained communication and irregular memory accesses. Many graph analysis problems feature such irregular memory and communication patterns, making them a good choice for performance evaluation. Graph500 is a popular benchmark built upon breadth-first search (BFS) on an undirected graph, which is well known for its sparse data accesses and fine-grained communication. In this research, we analyze and compare the scalability of multiple implementations of BFS using MPI-RMA against previously developed OpenSHMEM-based implementations optimized to maximize the benefits of one-sided communication. Additionally, we evaluate these implementations against the state-of-the-art MPI reference code using different numbers of processing elements and various problem sizes. Our experimental evaluation shows consistently improved performance with MPI-RMA over the best OpenSHMEM implementation on Graph500's BFS kernel with scales up to 32 nodes on the Pittsburgh Supercomputing Center Bridges-2 Regular Memory partition and University of Pittsburgh Center for Research Computing (Pitt CRC) MPI Cluster. Due to the nature of graph processing having a higher ratio of communication to computation, the communication latency hiding aspects of one-sided communication could not be fully exploited. While we demonstrate MPI-RMA to achieve~1.8x better performance over the MPI reference implementation on 32 nodes when only using 4 cores per node, the reference version was more performant in the majority of configurations tested. It is concluded that while one-sided communication has shown promising performance on some large-scale computing tasks, it remains difficult from a development standpoint to leverage the one-sided benefits on more complex kernels.

Read the paper · More papers on PaperTik