Design and Implementation of an IPC-based Collective MPI Library for Intel GPUs

Chen-Chun Chen, Goutham Kalikrishna Reddy Kuncham, Pouya Kousha, Hari Subramoni, Dhabaleswar K. DK Panda · 2024

With the rising demand for computing power in High-Performance Computing and Deep Learning applications, there is a noticeable trend in outfitting modern exascale clusters with accelerators. In recent years, Intel has been designing and developing GPU products and their associated ecosystems. Concurrently, application developers are transitioning their programs to Intel GPUs, seeking to maximize the computational capabilities of multi-GPU systems by utilizing efficient communication facilitated by modern GPU-aware MPI libraries. Hence, it is critical to design an efficient MPI collective library specifically tailored for Intel GPUs to optimize communication performance. In this paper, we proposed hybrid and IPC-based designs for data movement collective MPI operations on contemporary Intel GPU systems. For large message communication, we developed a comprehensive design for data movement collectives that surpasses reliance on basic send/recv pairs, effectively minimizing overheads. For small messages, we employ CPU staging techniques and compare various underlying libraries to ensure optimal performance. We evaluate the benefits of our designs at both the benchmark and application layers on the Intel DevCloud, utilizing 4 Intel GPUs connected with Xe Links. In benchmark-level evaluations, our Alltoall and Allgather implementations show a constant 100 µs improvement for large messages, while other operations like Bcast achieve a 72x performance enhancement compared to MPICH at 32MB. In application-level evaluations, our proposed designs demonstrate up to a 30% improvement for the HPC application heFFTe compared to the second-best solution using MPICH.

Read the paper · More papers on PaperTik