Cache oblivious parallelograms in iterative stencil computations

Robert Strzodka, Mohammed Shaheen, Dawid Pająk, Hans‐Peter Seidel · 2010

We present a new cache oblivious scheme for iterative stencil computations that performs beyond system bandwidth limitations as though gigabytes of data could reside in an enormous on-chip cache. We compare execution times for 2D and 3D spatial domains with up to 128 million double precision elements for constant and variable stencils against hand-optimized naive code and the automatic polyhedral parallelizer and locality optimizer PluTo and demonstrate the clear superiority of our results.

Read the paper · More papers on PaperTik