Inter-warp divergence aware execution on GPUs

Chulian Zhang · 2016

GPUs have appeared as very efficient many-core platform to execute applications with massive thread-level parallelism. GPUs achieve high throughput by running many threads concurrently and switching between them rapidly to hide memory latency. With the introduction of general-purpose programming models such as CUDA and OpenCL, many applications are motivated to use GPUs to accelerate compute intensive kernels and see impressive speedups. With the trend toward using GPUs for a diverse range of applications (e.g. vision and scientific computing) new challenges have been raised. New challenges have been raised for both algorithm and architecture designer when targeting GPUs for general-purpose applications.

Read the paper · More papers on PaperTik