Dynamic thread block launch

Jin Wang, Norman Rubin, Albert Sidelnik, Sudhakar Yalamanchili · 2015

GPUs have been proven effective for structured applications that map well to the rigid 1D-3D grid of threads in modern bulk synchronous parallel (BSP) programming languages. However, less success has been encountered in mapping data intensive irregular applications such as graph analytics, relational databases, and machine learning. Recently introduced nested device-side kernel launching functionality in the GPU is a step in the right direction, but still falls short of being able to effectively harness the GPUs performance potential.

Read the paper · More papers on PaperTik