Optimization strategy for a performance portable Vlasov code
Yuuichi Asahi, Guillaume Latu, Julien Bigot, V. Grandgirard · 2021
This paper presents optimization strategies applied on a kinetic plasma simulation code that makes use of Ope-nACC/OpenMP directives and Kokkos performance portable framework to run across multiple CPUs and GPUs. We evaluate the impacts of optimizations on multiple hardware platforms: Intel Xeon Skylake, Fujitsu Arm A64FX, and Nvidia Tesla P100 and V100. With vectorization and cache tuning, the Ope-nACC/OpenMP version achieved speedups of x1.07 to x 1.39. The Kokkos version in turn achieved speedups of x 1.00 to x 1.33 with Memory Layout and execution policy tuning. We obtain performance portability of multiple kernels ranging from 2.72% to 11.8% with OpenACC/OpenMP and 2.84% to 7.26% with Kokkos. Since the impact of optimizations under multiple combinations of kernels, devices and parallel implementations is demonstrated, this paper provides a widely available approach to accelerate codes keeping performance portability. To achieve good performance on both CPUs and GPUs, Kokkos could be a reasonable choice which offers more flexibility to manage multiple data and loop structures with a single codebase.