Multidisciplinary Simulation Acceleration using MultipleShared-Memory Graphical Processing Units - eScholarship
Jonathan Yashar Kemal, Roger L. Davis, John D. Owens · 2016
Multidisciplinary Simulation Acceleration using Multiple Shared-Memory Graphical Processing Units Jonathan Y. Kemal I , Roger L. Davis II , and John D. Owens III University of California, Davis, California, 95616 In this paper, we describe the strategies and programming techniques used in porting a multidisciplinary fluid/thermal interaction procedure to graphical processing units (GPUs). We discuss the strategies for selecting which disciplines or routines are chosen for use on GPUs rather than CPUs. In addition, we describe the programming techniques including use of Compute Unified Device Architecture (CUDA), mixed language (Fortran/C/CUDA) usage, Fortran/C memory mapping of arrays, and GPU optimization. We solve all equations using the multi-block, structured grid, finite-volume numerical technique, with the dual time-step scheme used for unsteady simulations. Our numerical solver code targets CUDA- capable Graphical Processing Units (GPUs) produced by NVIDIA. We use NVIDIA Tesla C2050/C2070 GPUs based on the Fermi architecture, and compare our resulting performance against Intel Xeon X5690 CPUs. Individual solver routines converted to CUDA typically run about 10 times faster on a GPU for sufficiently dense computational grids. We used a conjugate cylinder computational grid and ran a turbulent steady flow simulation using 4 increasingly dense computational grids. Our densest computational grid is divided into 13 blocks each containing 1033x1033 grid points, for a total of 13.87 million grid points or 1.07 million grid points per domain block. Comparing the performance of 8 GPUs to that of 8 CPUs, we obtain an overall speedup of about 6.0 when using our densest computational I Graduate student, Mechanical and Aerospace Department. Professor, Mechanical and Aerospace Engineering Department. III Professor, Electrical and Computer Engineering. II