Unraveling the Divergence of GPU Threads

Lucas John Vespa · 2018

We address the problem of performance issues caused by thread divergence due to branches in applications run on GPU. We eliminate all branches in GPU code through a method we call Algorithm Flattening (AF). AF is counterproductive for CPU code, but optimizes GPU code substantially. The seemingly counterproductive nature of AF is why this method is new and different from previous work. However, AF produces a multiple times speedup from an already GPU accelerated set of test algorithms. AF works by replacing branched code with arithmetic operations. AF increases ILP and processor utilization, and eliminates null operations due to branch divergence in SIMD devices. Although AF requires evaluation of code from all branch paths, this deterministic approach non-intuitively provides a substantial performance improvement for GPU applications, and allows GPUs to be used to accelerate previously unsuitable applications.

Read the paper · More papers on PaperTik