Summary of Investigations into Finite Volume Methods on GPUs
Laurence Kedward, Christian B Allen · AIAA SCITECH 2022 Forum · 2022
View Video Presentation: https://doi.org/10.2514/6.2022-0028.vid The trend for increased processor core numbers in both consumer and server CPUs, as well as the increasing use of many-core GPU accelerators for computation, has lead to a new paradigm of massive parallelism. GPUs in particular offer significant advantages in terms of memory bandwidth and floating-point performance as well as energy efficiency and price-to-performance. To leverage the new class of highly-parallel CPUs and GPUs, and close the gap between HPC capability and realised performance, numerical methods developed for the traditional multicore model must be reformulated for and combined with fine-grain thread-level parallelism. Previous work by the authors has investigated the importance of code structure and organisation when implementing finite volume methods on highly-parallel architectures like GPUs. In particular, the memory-bound nature of finite-volume kernels requires significant attention to be given to code structure and optimisation. Notably, the optimised code structure is very different from that of the original, with significant increases in performance being achieved through kernel merging. This paper summarises the key methodology and findings from that body of work including various code optimisations that were performed.