A Model for Weak Scaling to Many GPUs at the Basis of the Linpack Benchmark

David Rohr, Jan De Cuveland, V. Lindenstruth · 2016

Today, accelerator cards like GPUs are an important constituent of HPC clusters. For certain GPU-intense applications, the trend is shifting toward multi-GPU systems with four or more GPUs per compute node. This can increase the performance per dollar and the performance per watt. The Linpack benchmark is the standard tool for measuring the compute performance of supercomputers. Its standard implementation, HPL, cannot make use of GPUs on its own and other GPU-enabled versions of HPL show reduced efficiency on multi-GPU systems with four or more GPUs per node. In previous efforts we have developed a GPU-accelerated Linpack implementation in particular for AMD GPUs, which we employed on several clusters of a couple of hundred nodes each. Gustafson's law predicts perfect weak scaling with the number of processor cores, but it is not directly applicable to the number of GPUs in a system due to shared resources. In this paper we develop a model for general relations between the number of GPUs and the minimum problem size required to use them efficiently taking into account the limited PCIe bandwidth, memory bandwidth, and CPU resources. Based on this, we present an approach to scale our Linpack implementation to eight GPUs and achieve good GPU utilization in Linpack on future systems. Finally, we examine how energy efficiency, which has improved significantly with dual-and quad-GPU servers, scales to many-GPU systems.

Read the paper · More papers on PaperTik