Performance of Unstructured Finite Volume Code on a Cluster with Multiple GPUs per Node
Keith S. Obenschain, Andrew T. Corrigan, Gopal Patnaik · 49th AIAA Aerospace Sciences Meeting including the New Horizons Forum and Aerospace Exposition · 2011
This paper will investigate the performance of an unstructured finite volume code on a multi-CPU, multi-GPU cluster. This cluster attempts to balance IO, GPU, and CPU performance to accommodate a wide variety of codes. A new, unstructured finite volume code running in parallel using MPI/OpenMP and MPI/CUDA is presented. The performance of this code on a purpose-built GPU cluster is examined under a number of operating conditions. The GPU cluster is a collection of 24 compute nodes, each consisting of two, 6-core Intel Core i7 Processors and two NVIDIA GPUs with one to two QDR Infiniband ports connected to a switch. Eight of the compute nodes have two NVIDIA Fermi Tesla GPUs well connected with two QDR Infiniband cards. The remaining 18 have two NVIDIA Fermi Video GPUs. The use of multiple chipsets creates non-uniform access to both the GPUs and Infiniband, potentially creating bottlenecks when transferring data between the CPU and the GPU and between nodes. This paper will also explore these issues as well as potential solutions.