On-the-fly workload partitioning for integrated CPU/GPU architectures
Younghyun Cho, Florian Negele, Seohong Park, Bernhard Egger, Thomas R. Gross · 2018
Integrating CPUs and GPUs on the same die provides new opportunities for optimization, especially for irregular data-parallel workloads that fail to fully exploit the computational power of the GPU. Such workloads benefit from a proper partitioning between the CPU and the GPU. This paper presents an on-the-fly workload partitioning technique for irregular workloads on integrated architectures. Unlike existing work, no prior analysis of the workload is required. GPU kernels and input data are analyzed and optimized at runtime. The technique executes work chunks of similar load on the GPU and assigns irregular chunks to the CPU. Evaluated with various irregular workloads, the method achieves a 1.4x--7.1x speedup over GPU execution on AMD and Intel processors.