Enabling fast preemption via Dual-Kernel support on GPUs

Li-Wei Shieh, Kun-Chih Jimmy Chen, Hsueh-Chun Fu, Po-Han Peter Wang, Chia-Lin Yang · 2017

To consider QoS for resource-limited mobile systems, we introduce a fast preemption mechanism on GPUs. First, we involve a dual-kernel execution model to support fine-grained preemption, and a resource allocation policy to avoid resource fragmentation problem. Second, we propose a preemption victim selection scheme to reduce the throughput overhead while satisfying a required preemption latency. Evaluations show that we can reach very close to the ideal preemption scheme within 2% difference in terms of deadline violations. Furthermore, on average we improve GPU resource utilization by 2.93× over prior technique during preemption.

Read the paper · More papers on PaperTik