Compact Update Algorithm for Numerical Schemes with Cross Stencil for Data Access Locality

Andrey Vladimirovich Zakirov, Boris Azamatovich Korneev, Anastasia Yurievna Perepelkina · 2022

Accurate fluid simulations require high computing cost. 3D modelling of fluid dynamic field evolution on a discrete mesh takes large amount of data storage, and data access becomes performance bottleneck. Our work is concerned with the task of mitigating the limitations that are caused by finite memory throughput in the parallel simulations. We use LRnLA algorithms for this issue, where localized tasks combine updates on several time layers. In this paper, the compact update for DiamondTorre LRnLA algorithm is constructed. It further improves localization of DiamondTorre algorithm, which improves arithmetic intensity for cross-stencil schemes. The ratio of loaded data to fully updated data approaches 1. The compact update is implemented with CUDA C++ for a numerical scheme for the advection-diffusion equation. 50 GLU/sec (billion lattice updates per second) performance is obtained on Nvidia RTX3090, and the maximal performance of almost 300 GLU/sec is obtained on an 8 GPU workstation. Note that the main data storage is in CPU RAM memory, but the host-device data exchange is concealed by temporal blocking: with appropriate the data transfers are concealed by the computing operations and do not affect the performance.

Read the paper · More papers on PaperTik