A Performance Evaluation of Dynamic Parallelism for Fine-Grained, Irregular Workloads
Max Plauth, Frank Feinbube, Frank Schlegel, Andreas Polze · International Journal of Networking and Computing · 2016
GPU compute devices have become very popular for general purpose computations. However, the SIMD-like hardware of graphics processors is currently not well suited for irregular workloads, like searching unbalanced trees. In order to mitigate this drawback, NVIDIA introduced an extension to GPU programming models called Dynamic Parallelism. This extension enables GPU programs to spawn new units of work directly on the GPU, allowing the refinement of subsequent work items based on intermediate results without any involvement of the main CPU.