Autotuning Structured Grid Kernels

Kaushik Datta, Samuel Williams, Vasily M. Volkov, Ashim Kumar Datta, Mark Murphy, Vasily M. Volkov, Samuel Williams, Jonathan T. Carter, Leonid Oliker, John M. Shalf, Katherine Yelick · 2008

Code Generators •Kernel-specific •Perl script generates 1000’s of code variations •Autotuner searches over all possible implementations (sometimes guided by a performance model to prune the space) to find the optimal configuration •Optimizations included in this work: • NUMA-Aware collocates data with the threads processing it • Array Padding avoids conflicts in the L1/L2 • Thread/Cache minimizes cache misses and memory traffic Blocking • Vectorization avoids rolling the TLB • Unrolling/DLP compensates for poor compilers • SW Prefetching attempts to hide L2 and DRAM latency • SIMDization compensates for poor compilers, and streaming stores minimize memory traffic Autotuning Lattice Methods

Read the paper · More papers on PaperTik