Enabling collaborative heterogeneous computing

Yifan Sun · 2020

GPUs have been accelerating a wide range of algorithms and applications with their massively parallel computing capabilities. Today, as GPU programs work under the close supervision of a CPU, computing platforms are usually equipped with both CPUs and GPUs. Most existing GPU programs underutilize the computing capabilities of the CPU since the CPU is frequently idle while GPUs are executing. Additionally, existing work primarily focuses on using a single GPU, thus failing to utilize the computing power of multi-GPU systems. In this dissertation, we propose a new heterogeneous computing paradigm--Collaborative Heterogeneous Computing. Collaborative Heterogeneous Computing leverages fine-grained CPU-GPU communication mechanisms in computation with the use of CPUs and multiple GPUs simultaneously. Collaborative Heterogeneous Computing can be divided into CPU-GPU collaborative computing and Multi-GPU collaborative computing. For CPU-GPU collaborative computing, we first categorize 7 CPU-GPU collaborative execution patterns. With the guidance of the summarized CPU-GPU execution patterns, we implement Hetero-Mark, a benchmark suite that enables users to explore CPU-GPU collaborative execution patterns. We also design and develop Multi2Sim-HSA, an HSA system emulator that emulates behaviors of CPUs and GPUs in CPU-GPU collaborative execution applications. In addition, we observe that CPU-GPU communication over the PCIe network causes congestion and can become a major performance bottleneck. To address this issue, we propose a priority-based PCIe scheduling algorithm that can improve device utilization and satisfy Quality of Service (QoS) requirements. For multi-GPU collaborative computing, we develop MGPUMark, a benchmark suite that explores multi-GPU collaborative execution and fine-grained inter-GPU communication. Benchmarks in MGPUMark enable exploration of inter-GPU communication patterns. We also design and develop MGPUSim, a high-performance, high-flexibility, and high-accuracy GPU simulator. MGPUSim can faithfully model behaviors of GPUs in multi-GPU collaborative computing. We identify that the main performance bottlenecks occur due to long-latency, low-bandwidth, inter-GPU interconnects. Thus, reducing inter-GPU communication traffic can improve multi-GPU system performance. We develop the Locality APIs, a set of GPU programming APIs that enables programmers to place the computing threads and associated data on a multi-GPU system to reduce inter-GPU traffic. We also introduce Progressive Page-Splitting Migration (PASI), a hardware-based solution that allows GPUs to adjust data placement to avoid inter-GPU traffic.--Author's abstract

Read the paper · More papers on PaperTik