Optimizing All-to-All and Allgather Communications on GPGPU Clusters
Ashish Kumar Singh · OhioLink ETD Center (Ohio Library and Information Network) · 2012
High Performance Computing (HPC) is rapidly becoming an integral part of Science, Engineering and Business.Scientists and engineers are leveraging HPC solutions to run their applications that require high bandwidth, low latency, and very high compute capabilities.General Purpose Graphics Processing Units (GPGPUs) are becoming more popular within the HPC community because of their highly parallel structure, which makes it possible for applications to gain multi-x performance gain.The Tianhe-1A and Tsubame systems received significant attention for their architectures that leverage GPGPUs.Increasingly many scientific applications that were originally written for CPUs using MPI for parallelism are being ported to these hybrid CPU-GPU clusters.In the traditional sense, CPUs perform computation while the MPI library takes care of communication.When computation is performed on GPGPUs, the data has to be moved from device memory to main memory before it can be used in communication.Though GPGPUs provide huge compute potential, the data movement to and from GPGPUs is both a performance and productivity bottleneck.Recently, the MVAPICH2 MPI library has been modified to directly support point-to-point MPI communication from the GPU memory [33].Using this support, programmers do not need to explicitly move data to main memory before using MPI.This feature also enables performance improvement due to tight integration of GPU data movement and MPI internal protocols.