Automatic Code Generation and Adaptive Grid Scheduling for GPU Cluster Computing
Xuyuan Jin · 2012
Recent advances in GPUs (graphics processing units) lead to mas-sively parallel hardware that is easily programmable and widely ap-plied in areas which require intensive computation besides graph-ics acceleration. The appearance of GPU clusters gains popularity in the scientific computing community, and the study on GPU clus-ters becomes an increasingly hot issue. While extending a single-GPU system to a multi-GPU cluster, the workload is spawned onto a number of GPU devices, and device-level and machine-level par-allelism are achieved on top of the native SIMD thread-level paral-lelism. This requires a programming model extension and a work-load scheduling strategy, but currently most of the programmers have to perform the extension manually and schedule the work-loads in a naive way. Analytical performance models of GPUs exist, however, prediction of the performance of a GPU cluster is more complicated and no dedicated study has been done on this topic. In this paper, programming model is extended to fit the GPU clus-ter and propose a tool for automatic code generation and adaptive workload scheduling with which the application gains a 3x speedup from the cluster compared with executed on a single GPU. Besides, results show that with an extended analytical performance model, the performance of a GPU cluster is predictable.