Optimisation of a Finite-Volume Test-bench Code for Highly Parallel Architectures
Laurence Kedward, Christian B Allen · AIAA Scitech 2021 Forum · 2021
View Video Presentation: https://doi.org/10.2514/6.2021-0143.vid Previous work by the authors investigated code structure and optimisation for a finite-volume test-bench code on highly parallel computing architectures such as GPUs. Preliminary results show an overall speedup of between 18x and 24x on two high performance consumer GPUs compared to 16 threads of a ninth generation Intel CPU. Further results are presented here with a focus on the highly memory-bound nature of typical finite-volume kernels. Two methods are presented for improving performance under memory bandwidth limitations: kernel merging with shared memory use, and a mixed-precision implicit solver. The kernel merging code optimisation, which exploits fast on-chip shared memory in GPU architectures, increases arithmetic intensity and reduces kernel launch overheads. Significant performance improvements are seen on both GPU architectures as well as on an Intel CPU due to better cache use. Results presented for a simple mixed-precision implicit solver verify that the linear system can be solved in single-precision while retaining full double-precision for the non-linear flow residual. Moreover, close to a 2x speedup is achieved on the GPU architectures when using single-precision for the linear system due to the reduced memory bandwidth requirements.