Cache Blocking Strategies Applied to Flux Reconstruction

Semih Akkurt · 2021

On modern hardware architectures, the performance of Flux Reconstruction (FR) methods for tensor product elements can be limited by memory bandwidth. In general, these methods are implemented as a chain of distinct kernels. Often, a dataset which has just been written to main memory by a kernel is read back immediately by the next kernel. The accepted solution for such a redundant expenditure of memory bandwidth is kernel fusion. However, on a practical level kernel fusion requires that the source for all kernels be available, thus preventing calls to certain third-party library functions. Moreover, it can add substantial complexity to a codebase. An alternative to full kernel fusion is using cache blocking strategies, which has become practically possible in recent years due to growth in the size of CPU L2 cache. In this approach, kernels remain distinct, and are executed one after another on small chunks of data that can fit in the cache, as opposed to on full datasets. These chunks of data stay in the cache and whenever a kernel requests access to data that is already in the cache, memory bandwidth is saved. In this talk, a variety of kernel grouping configurations will be presented for FR based Euler and Navier-Stokes solvers, alongside associated theoretical memory bandwidth savings, and actual achieved performance gains when implemented in the PyFR solver. Inviscid and viscous Taylor-Green Vortex test cases are used as benchmark for Euler and Navier-Stokes solvers, and the most performant strategies lead to a speedup of approximately 2.63x and 3.01x, respectively.

Read the paper · More papers on PaperTik