Performance and Programmability of MPI+X Integration with CUDA, HIP, SYCL, OpenACC, and OpenMP Offloading for Supercomputing: A Case Study on Dense Matrix–Vector Multiplication

Ezhilmathi Krishnasamy, James D. Trotter, Xing Hui Cai, Dirk Pleiter, L. Kos, Laura Saavedra, Pascal Bouvry · 2026

Dense matrix-vector multiplication is a fundamental operation widely used in various application areas, including deep learning, computer graphics, and numerical analysis. With the increasing prevalence of large, heterogeneous high-performance computing (HPC) systems equipped with multiple Graphical Processing Units (GPUs), it is crucial to analyze effective strategies for utilizing these accelerators to enhance the efficiency of numerical computation. This study thus centers on the intricacies of accelerating dense matrix-vector multiplication on multi-GPU systems. We explore the combination of the Message Passing Interface (MPI) with various programming models—collectively referred to as MPI + X—highlighting its potential for advancing capabilities in large HPC environments, particularly those aimed at exascale computing and beyond. Here, "X" encompasses a range of programming models compatible with MPI. Specifically, this study investigates Compute Unified Device Architecture (CUDA), HIP, SYCL, Open Accelerators (OpenACC), and Open Multi-Processing (OpenMP) Offloading programming models within the context of large heterogeneous systems, as all can be integrated with MPI. We conduct a thorough examination of parallel dense matrix-vector multiplication and provide a detailed performance analysis of these programming models on NVIDIA, AMD, and Intel GPUs, addressing scenarios involving both GPU-aware and non-GPU-aware MPI. Our findings are substantiated by openly available code, which is essential for validation and the progression of scientific research.

Read the paper · More papers on PaperTik