Communication Profiling and Characterization of Deep Learning Workloads on Clusters with High-Performance Interconnects

Ammar Ahmad Awan, Arpan Jain, Ching-Hsiang Chu, Hari Subramoni, Dhableswar K. Panda · 2019

Heterogeneous HPC systems with GPUs are increasingly getting equipped with on-node interconnects like PCIe and NVLink and inter-node interconnects like InfiniBand and Omni-Path. However, the efficient exploitation of these interconnects brings forth many challenges for MPI+CUDA applications. Little exists in the literature that captures the impact of these interconnects on emerging application areas like distributed Deep Learning (DL). In this paper, we choose Horovod; a distributed training middleware, to analyze and profile high-level application workloads (e.g., Training ResNet-50) instead of MPI microbenchmarks. It is challenging to use existing profilers like mpiP and nvprof as they only offer a black box approach and cannot profile emerging communication libraries like NCCL. To address this, we developed a profiler for Horovod that enables profiling of various communication primitives including MPI_Allreduce and ncclAllreduce for gradient exchange as well for Horovod's communication threads and response caches. We analyze the following metrics to gain insights into network-level performance on different interconnects: 1) Message size with tensor fusion, 2) Message size without tensor fusion, 3) Number of MPI and NCCL calls made for each message size, and 4) Time taken by each NCCL and/or MPI call. We also correlate these low-level statistics to higher level end-to-end training metrics like images per second. Three keys insights we gained are: 1) Horovod tensor fusion offers slight performance gains (up to 5%) for CPU-based training on InfiniBand systems, 2) For GPU-based training, disabling tensor fusion improved performance (up to 17%) for GPUs connected with PCIe, and 3) The allreduce latency profiles show some extreme performance variations for non-power-of-two message sizes for both CPUs and GPU on all interconnects when tensor fusion is enabled. To provide a comprehensive view of performance, we use a wide variety of systems with CPUs like Intel Skylake, AMD EPYC, and IBM POWER9, GPUs like Volta V100, and interconnects like PCIe, NVLink, InfiniBand, and Omni-Path.

Read the paper · More papers on PaperTik