Model-Based Warp Overlapped Tiling for Image Processing Programs on GPUs

Abhinav Jangda, Arjun Guha · 2020

Domain-specific languages that execute image processing pipelines on GPUs, such as Halide and Forma, operate by 1)~dividing the image into overlapped tiles, and 2)~fusing loops to improve memory locality. However, current approaches have limitations: 1)~they require intra thread block synchronization, which has a nontrivial cost, 2)~they must choose between small tiles that require more overlapped computations or large tiles that increase shared memory access (and lowers occupancy), and 3) their autoscheduling algorithms use simplified GPU models that can result in inefficient global memory accesses.

Read the paper · More papers on PaperTik