Launch-Time Optimization of OpenCL GPU Kernels

Andrew S. D. Lee, Tarek S. Abdelrahman · 2017

OpenCL compiles a GPU kernel first and then launches it for execution, providing the kernel at this launch with its arguments and its launch geometry. Although some of the kernel inputs and the launch geometry remain constant across all threads during execution, the compiler is unable to treat them as such, which limits its ability to apply several optimizations, including constant propagation, constant folding, strength reduction and loop unrolling.

Read the paper · More papers on PaperTik