Porting of FEFLO to Multi-GPU Clusters
Andrew T. Corrigan, Rainald Löhner · 49th AIAA Aerospace Sciences Meeting including the New Horizons Forum and Aerospace Exposition · 2011
FEFLO is an adaptive, edge-based finite element code for the solution of compressible and incompressible flow. It is primarily written in Fortran 77 and has been ported to vector, shared memory parallel and distributed memory parallel machines. FEFLO is currently undergoing the process of being ported to run on graphics hardware (GPUs), using semiautomatic techniques. In previous work, a GPU version of FEFLO was presented, in which, for many run configurations, all mesh-sized loops required throughout time-stepping were ported. This approach simultaneously achieves the fine-grained parallelism required to fully exploit the capabilities of many-core GPUs, completely avoids the crippling bottleneck of GPU-CPU data transfer, and uses a transposed memory layout to meet the distinct memory access requirements posed by GPUs. The present work describes the next step of this porting effort, namely to integrate GPU-based, fine-grained parallelism with MPIbased, coarse-grained parallelism, in order to achieve a code capable of running on multiGPU clusters. This is done in a semi-automated fashion: the existing Fortran-MPI code is preserved, with the translator inserting data transfer calls as required. Performance benchmarks indicate up to a factor of two performance advantage of the NVIDIA Tesla M2050 GPU over the six-core Intel Xeon X5670 CPU, for certain run configurations. In addition, good scalability is observed when running across multiple GPUs.