SCALABLE - D5.3 Efficient architecture-specific compute kernels
Philipp Suffa, Markus Holzer, Ulrich Rüde, Harald Köstler · Zenodo (CERN European Organization for Nuclear Research) · 2023
The goal of tasks 5.1 and 5.2 was to extend the code generation pipeline of lbmpy to support a full CFD simulation, which can be run with sparse data kernels. That means that we integrated sparse LBM kernels as well as sparse boundary kernels and communication kernels into the generation pipeline. The goal of task 5.3 is to optimize these sparse kernels to achieve better performance results on CPUs as well as on GPUs. Therefore, a description of the automatic code generation for architecture-specific sparse data kernels will be given in chapter 2. The first optimization for the sparse data kernels is an in-place streaming pattern. The implementation of an in-place streaming pattern, here the AA-pattern, reduces the amount of memory needed for the simulation and, more importantly, reduces the number of memory accesses and thus increases the performance for the LBM on CPUs, as it is shown in chapter 3. Furthermore, in chapter 4, we utilized communication hiding on GPUs to achieve greater scaling efficiency on this hardware. Lastly, we demonstrate the increase in performance of the optimized generated sparse kernels. Therefore, we integrate the generated sparse kernels in our massive parallel multi-physics framework waLBerla to enable parallel execution with great scalability. Then we compare the sparse kernels with and without optimizations by running a turbulent channel scenario, which is one of the target applications of the SCALABLE project. These scaling runs are made on the CPU cluster Juwels-Cluster as well as on the GPU cluster Juwels-Booster.