Exploiting reuse and vectorization in blocked stencil computations on CPUs and GPUs
Tuowen Zhao, Protonu Basu, Samuel Williams, Mary W. Hall, Hans Johansen · 2019
Stencil computations in real-world scientific applications may contain multiple interrelated stencils, have multiple input grids, and use higher order discretizations with high arithmetic intensity and complex expression structures. In combination, these properties place immense demands on the memory hierarchy that limit performance. Blocking techniques like tiling are used to exploit reuse in caches. Additional fine-grain data blocking can also reduce TLB, hardware prefetch, and cache pressure.