Characterizing numascale clusters with GPUs: MPI-based and GPU interconnect benchmarks

Malik Muhammad Zaki Murtaza Khan, Anne Cathrine Elster · 2016

Modern HPC clusters are increasingly heterogeneous both in processor types, topologies of computing, communication and storage resources. In this paper, we describe how to use benchmarking, to characterize the high-speed interconnect, NumaConnect, associated with a shared-memory Numascale cluster system with GPUs, constituting a novel testbed at NTNU. Numascale systems include a unique node controller, NumaConnect, based on the FPGA or ASIC-based NumaChip, depending on system vendor requirements. The system's interconnects uses AMD's HyperTransport protocol, and provide a cache-coherent shared-memory single image operating system. Our system has, in addition, a GPU added to each server blade. Our characterizations efforts target the NumaConnect which includes an RDMA-type Block Transfer Engine (BTE). The BTE is used by Byte Transfer Libraries such as the NumaConnect BTL (NC-BTL) for message passing (MPI) or BLACS. To characterize our Numascale system, we use several benchmark suites including: our own SimpleBench that includes ping-pong, MPI-Reduce and MPI-Barrier tests; two well-known MPI benchmark suites: the NAS Parallel Benchmarks (NPB)-MPI, the OSU microbenchmarks; as well as Nvidia's Bandwidth test for GPUs. Our results show that it is generally very beneficial to use MPI or other libraries that use the NC-BTL library. In fact, on selected OSU and NPB benchmarks, we achieve order-of-magnitude performance improvements on communication and synchronization costs on these benchmarks when using NC-BTL.

Read the paper · More papers on PaperTik