Architectural and runtime enhancements for dynamically controlled multi-level concurrency on GPUs
Yash Ukidave · 2015
GPUs have gained tremendous popularity as accelerators for a broad class of applications belonging to a number of important computing domains. Many applications have achieved significant performance gains using the inherent parallelism offered by GPU architectures. Given the growing impact of GPU computing, there is a growing need to provide improved utilization of compute resources and increased application throughput. Modern GPUs support concurrent execution of kernels from a single application context in order to increase the resource utilization. However, this support is limited to statically assigning compute resources to multiple kernels, and lacks the flexibility to adapt resources dynamically. The degree of concurrency present in a single application may also be insufficient to fully exploit the resources on a GPU.